Datasets ডেটাসেট
63 Bangla datasets across 26 tasks. Select a dataset name for details and a citation.
63 shown
| Dataset | Task | Size | License | Year | Link |
|---|---|---|---|---|---|
| Machine Translation | 4,653 sentences | CC BY 4.0 | 2026 | ↗✓ 2026-07-26 | |
Regional sentences from Rangpur, Barisal, Narail, and Khulna with standard Bangla and English translations. Source: BanglaRegionalTextCorpus (Data in Brief 2026). @article{Ahmed_2026, title={BanglaRegionalTextCorpus: A curated dataset for four regional bangla dialects with standard Bangla and English translation}, volume={65}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2026.112585}, DOI={10.1016/j.dib.2026.112585}, journal={Data in Brief}, publisher={Elsevier BV}, author={Ahmed, Md. Tofael and Koli, Zannatul Mawa and Rahat, Azmain Mahtab and Akhter, Taslima and Ayman, Umme}, year={2026}, month=Apr, pages={112585} } | |||||
| LLM Benchmarks | 1,719 instances | CC BY 4.0 | 2026 | ↗✓ 2026-08-23 | |
Sociopragmatic LLM benchmark over Bangla address terms, kinship reasoning, and social customs. Source: BanglaSocialBench (ACL SRW 2026). @inproceedings{sijan-etal-2026-banglasocialbench,
title = "{B}angla{S}ocial{B}ench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of {LLM}s in Bangladeshi Social Interaction",
author = "Sijan, Tanvir Ahmed and
Rifat, S. M. Golam and
Partha, Pankaj and
Islam, Md. Tanjeed and
Anwar, Md Musfique",
editor = "T.Y.S.S., Santosh and
Rodriguez, Juan Diego and
de Gibert, Ona",
booktitle = "Proceedings of the 64th Annual Meeting of the {A}ssociation for {C}omputational {L}inguistics (Volume 4: Student Research Workshop)",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.acl-srw.22/",
doi = "10.18653/v1/2026.acl-srw.22",
pages = "247--280",
ISBN = "979-8-89176-393-7"
} | |||||
| ASR / TTS | 191 recordings / 158.6 hrs | Mixed | 2026 | ↗✓ 2026-08-10 | |
Long-form Bangla ASR corpus of 792k words from 11 YouTube channels, built by subtitle extraction with human-in-the-loop transcript verification. Source: Bengali-Loop (DL Sprint 4.0, Tabib et al.). @misc{tabib2026bengaliloopcommunitybenchmarkslongform,
title={Bengali-Loop: Community Benchmarks for Long-Form Bangla ASR and Speaker Diarization},
author={H. M. Shadman Tabib and Istiak Ahmmed Rifti and Abdullah Muhammed Amimul Ehsan and Somik Dasgupta and Md Zim Mim Siddiqee Sowdha and Abrar Jahin Sarker and Md. Rafiul Islam Nijamy and Tanvir Hossain and Mst. Metaly Khatun and Munzer Mahmood and Rakesh Debnath and Gourab Biswas and Asif Karim and Wahid Al Azad Navid and Masnoon Muztahid and Fuad Ahmed Udoy and Shahad Shahriar Rahman and Md. Tashdiqur Rahman Shifat and Most. Sonia Khatun and Mushfiqur Rahman and Md. Miraj Hasan and Anik Saha and Mohammad Ninad Mahmud Nobo and Soumik Bhattacharjee and Tusher Bhomik and Ahmmad Nur Swapnil and Shahriar Kabir},
year={2026},
eprint={2602.14291},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2602.14291},
} | |||||
| Text Classification | 10,283 messages | CC BY 4.0 | 2026 | ↗✓ 2026-07-24 | |
Tri-class Bangla SMS/Telegram corpus for spam, ham, and promotional message classification. Source: Mendeley Data (2026). @misc{https://doi.org/10.17632/5wrm959d6f.1,
doi = {10.17632/5WRM959D6F.1},
url = {https://data.mendeley.com/datasets/5wrm959d6f/1},
author = {Tahmid, Labib and Abedin, Istiak and Ahmed, Kawsar and Hassan Chowdhury , Zinedin and Sultana, Babe },
keywords = {Computer Science, Artificial Intelligence, Cybersecurity, Natural Language Processing, Machine Learning},
title = {BTTC - A Bangla Tri-class Text Corpus for Spam, Ham, and Promotional Messages},
publisher = {Mendeley Data},
year = {2026}
} | |||||
| Machine Translation | 17,300 sentence pairs | CC BY 4.0 | 2026 | ↗✓ 2026-07-26 | |
Parallel Chittagonian-Bangla and Sylheti-Bangla corpora with 8,000 and 9,300 sentence pairs. Source: Kothon (Data in Brief 2026). @article{Faisal_2026, title={Kothon: A large-scale dataset for machine translation of the Chittagonian and Sylheti dialects into standard Bangla}, volume={66}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2026.112789}, DOI={10.1016/j.dib.2026.112789}, journal={Data in Brief}, publisher={Elsevier BV}, author={Faisal, Md. Atique and Sadaf, Farhan and Chowdhury, Dipta and Azrof, H.M. and Tanmay, Monojit Paul}, year={2026}, month=June, pages={112789} } | |||||
| ASR / TTS | 882 hrs | CC BY 4.0 | 2026 | ↗✓ 2026-07-22 | |
Long-form multi-speaker Bangla ASR and speaker-diarization corpus with aligned transcriptions and timestamps. Gated access. Source: Lipi-Ghor-882 (DL Sprint 4.0, Team Villagers). @misc{hasan2026makehardheareasy,
title={Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment},
author={Sanjid Hasan and Risalat Labib and A H M Fuad and Bayazid Hasan},
year={2026},
eprint={2602.23070},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2602.23070},
} | |||||
| Question Answering | 87,805 QA pairs | CC BY-NC 4.0 | 2026 | ↗✓ 2026-08-23 | |
Educational QA from 50 NCTB textbooks with answerable, unanswerable, and adversarial instances. Source: NCTB-QA (arXiv 2026). @misc{eyasir2026nctbqalargescalebanglaeducational,
title={NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance},
author={Abrar Eyasir and Tahsin Ahmed and Muhammad Ibrahim},
year={2026},
eprint={2603.05462},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2603.05462}
} | |||||
| Hate Speech Detection | 4,027 texts | CC BY 4.0 | 2025 | ↗✓ 2026-07-26 | |
Bangla religious-aggression texts with English translations, labelled hate speech, vandalism, atrocity, or no aggression. Source: ALERT (Data in Brief 2025). @article{Rashid_2025, title={ALERT: A benchmark Bengali dataset for identifying and categorizing religiously aggressive texts}, volume={63}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2025.112094}, DOI={10.1016/j.dib.2025.112094}, journal={Data in Brief}, publisher={Elsevier BV}, author={Rashid, Suhana Binta and Piyas, Bibhas Roy Chowdhury and Rahman, Sadia and Preenon, Bijoy Roy Chowdhury}, year={2025}, month=Dec, pages={112094} } | |||||
| Named Entity Recognition | 17,405 sentences | CC BY 4.0 | 2025 | ↗✓ 2026-07-19 | |
Bangla regional NER across five dialects with 10 entity types, BIO-tagged. Source: ANCHOLIK-NER (2025). @misc{paul2025ancholiknerbenchmarkdatasetbangla,
title={ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition},
author={Bidyarthi Paul and Faika Fairuj Preotee and Shuvashis Sarker and Shamim Rahim Refat and Shifat Islam and Tashreef Muhammad and Mohammad Ashraful Hoque and Shahriar Manzoor},
year={2025},
eprint={2502.11198},
archivePrefix={arXiv},
primaryClass={cs.CL},
doi={https://doi.org/10.1371/journal.pone.0342786},
url={https://arxiv.org/abs/2502.11198},
} | |||||
| LLM Benchmarks | 13,497 questions | MIT | 2025 | ↗✓ 2026-07-19 | |
Bengali multiple-choice evaluation suite over 50 subjects at four difficulty levels, plus the harder B-REASO HEAVY subset. Source: B-REASO (Findings EMNLP 2025). @inproceedings{hosain-morol-2025-b,
title = {{B}-{REASO}: A Multi-Level Multi-Faceted {B}engali Evaluation Suite for Foundation Models},
author = {Hosain, Md Tanzib and Morol, Md Kishor},
booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2025},
year = {2025},
pages = {9260--9274},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2025.findings-emnlp.492/},
doi = {10.18653/v1/2025.findings-emnlp.492}
} | |||||
| Sentiment Analysis | 15,860 instances | CC BY-NC-SA 4.0 | 2025 | ↗✓ 2026-07-26 | |
Aspect-based sentiment corpus across 21 domains with manually annotated aspect terms and positive, neutral, or negative labels. Source: BABSA (Data in Brief 2026). @article{Siam_2026, title={BABSA: A large scale bangla aspect based sentiment analysis dataset}, volume={65}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2026.112604}, DOI={10.1016/j.dib.2026.112604}, journal={Data in Brief}, publisher={Elsevier BV}, author={Siam, M.F.R. and Wafa, Narmeen}, year={2026}, month=Apr, pages={112604} } | |||||
| Text Classification | 60,000 articles | Apache 2.0 | 2025 | ↗✓ 2026-07-26 | |
Fake-news corpus with 47,000 authentic and 13,000 manually annotated fake articles across 13 categories. Source: BanFakeNews-2.0 (IndoNLP 2025). @inproceedings{shibu-etal-2025-scarcity,
title = "From Scarcity to Capability: Empowering Fake News Detection in Low-Resource Languages with {LLM}s",
author = "Shibu, Hrithik Majumdar and
Datta, Shrestha and
Miah, Md. Sumon and
Sami, Nasrullah and
Chowdhury, Mahruba Sharmin and
Islam, Md Saiful",
editor = "Weerasinghe, Ruvan and
Anuradha, Isuri and
Sumanathilaka, Deshan",
booktitle = "Proceedings of the First Workshop on Natural Language Processing for Indo-Aryan and Dravidian Languages",
month = jan,
year = "2025",
address = "Abu Dhabi",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.indonlp-1.12/",
pages = "100--107"
} | |||||
| Named Entity Recognition | 2,980 texts | CC BY 4.0 | 2025 | ↗✓ 2026-07-26 | |
Manually annotated medical text covering six entity types, with a separate English translation. Source: Bangla-MedER (Data in Brief 2026). @article{Sheikh_2026, title={Bangla-MedER: An annotated Bangla dataset for multi-type medical entity recognition from medical text}, volume={66}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2026.112705}, DOI={10.1016/j.dib.2026.112705}, journal={Data in Brief}, publisher={Elsevier BV}, author={Sheikh, Rubel and Rafiq, Shifat Ara and Hasan, Md. Mehedi and Ahmed, Shakil and Aurpa, Tanjim Taharat and Akter, Farzana}, year={2026}, month=June, pages={112705} } | |||||
| Hate Speech Detection | 1,004 comments | CC BY 4.0 | 2025 | ↗✓ 2026-07-26 | |
Context-aware toxic-comment corpus retaining news metadata and neighbouring comments, with balanced toxic/non-toxic labels. Source: Bangla-ToCo (Data in Brief 2025). @article{Rupa_2025, title={Bangla-ToCo: A context-aware dataset for Bengali toxic comment detection}, volume={63}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2025.112277}, DOI={10.1016/j.dib.2025.112277}, journal={Data in Brief}, publisher={Elsevier BV}, author={Rupa, Sayma Akter and Anwar, Md. Musfique and Ritu, Nadia Afrin and Arifin, Khandoker Nosiba}, year={2025}, month=Dec, pages={112277} } | |||||
| LLM Benchmarks | ~1,700 problems | Research only | 2025 | ↗✓ 2026-07-19 | |
Bangla grade 6-8 math word problems for testing LLM mathematical reasoning. Source: BanglaMATH (MathNLP 2025). @inproceedings{prama-etal-2025-banglamath,
title = "{B}angla{MATH} : A {B}angla benchmark dataset for testing {LLM} mathematical reasoning at grades 6, 7, and 8",
author = "Prama, Tabia Tanzin and
Danforth, Christopher M. and
Dodds, Peter",
editor = "Valentino, Marco and
Ferreira, Deborah and
Thayaparan, Mokanarangan and
Ranaldi, Leonardo and
Freitas, Andre",
booktitle = "Proceedings of The 3rd Workshop on Mathematical Natural Language Processing (MathNLP 2025)",
month = nov,
year = "2025",
address = "Suzhou, China",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.mathnlp-main.10/",
doi = "10.18653/v1/2025.mathnlp-main.10",
pages = "134--149",
ISBN = "979-8-89176-348-7"
} | |||||
| Sentiment Analysis | 12,089 comments | CC BY 4.0 | 2025 | ↗✓ 2026-07-26 | |
Balanced ternary Facebook-comment corpus labelled sarcastic, non-sarcastic, or neutral. Source: BanglaSarc3 (Data in Brief 2025). @article{Biswas_2025, title={BanglaSarc3: A benchmark dataset for Bangla sarcasm detection from social media to advance Bangla NLP}, volume={62}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2025.111953}, DOI={10.1016/j.dib.2025.111953}, journal={Data in Brief}, publisher={Elsevier BV}, author={Biswas, Susmoy and Zahid, Md. Mostafizur Rahman and Rabeya, Mst Taposi and Abedin, Md. Minhazul and Bijoy, Md. Hasan Imam and Rahman, Md. Sadekur}, year={2025}, month=Oct, pages={111953} } | |||||
| Hate Speech Detection | 19,203 comments | MIT | 2025 | ↗✓ 2026-08-23 | |
YouTube comments annotated for binary hate, hate-category, and target-group labels. Source: BanHate (BLP 2025). @inproceedings{raquib-etal-2025-banhate,
title = "{B}an{H}ate: An Up-to-Date and Fine-Grained {B}angla Hate Speech Dataset",
author = "Raquib, Faisal Hossain and
Mazumder, Akm Moshiur Rahman and
Fuad, Md Tahmid Hasan and
Ishmam, Md Farhan and
Fahim, Md",
editor = "Alam, Firoj and
Kar, Sudipta and
Chowdhury, Shammur Absar and
Hassan, Naeemul and
Hoque, Enamul and
Laskar, Md Tahmid Rahman and
Mohiuddin, Tasnim and
Rony, Md Rashad Al Hasan",
booktitle = "Proceedings of the Second Workshop on Bangla Language Processing (BLP-2025)",
month = dec,
year = "2025",
address = "Mumbai, India",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.banglalp-1.19/",
doi = "10.18653/v1/2025.banglalp-1.19",
pages = "237--248",
ISBN = "979-8-89176-314-2"
} | |||||
| Named Entity Recognition | 85,175 sentences | CC BY-NC-ND 4.0 | 2025 | ↗✓ 2026-08-23 | |
Human-annotated Bangla NER corpus spanning 29 domains and 10 entity classes. Source: BanNERD (Findings NAACL 2025). @inproceedings{mahtab-etal-2025-bannerd,
title = "{B}an{NERD}: A Benchmark Dataset and Context-Driven Approach for {B}angla Named Entity Recognition",
author = "Mahtab, Md. Motahar and
Khan, Faisal Ahamed and
Islam, Md. Ekramul and
Chowdhury, Md. Shahad Mahmud and
Chowdhury, Labib Imam and
Afrin, Sadia and
Ali, Hazrat and
Rashid, Mohammad Mamun Or and
Mohammed, Nabeel and
Amin, Mohammad Ruhul",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2025",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-naacl.380/",
doi = "10.18653/v1/2025.findings-naacl.380",
pages = "6822--6843",
ISBN = "979-8-89176-195-7"
} | |||||
| Hate Speech Detection | 37,300 samples | MIT | 2025 | ↗✓ 2026-07-19 | |
Multi-label transliterated (romanized) Bangla hate speech dataset from YouTube across 7 target classes. Source: BanTH (Findings NAACL 2025). @inproceedings{haider-etal-2025-banth,
title = "{B}an{TH}: A Multi-label Hate Speech Detection Dataset for Transliterated {B}angla",
author = "Haider, Fabiha and
Shifat, Fariha Tanjim and
Ishmam, Md Farhan and
Sourove, Md Sakib Ul Rahman and
Barua, Deeparghya Dutta and
Fahim, Md and
Bhuiyan, Md Farhad Alam",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2025",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-naacl.403/",
doi = "10.18653/v1/2025.findings-naacl.403",
pages = "7232--7251",
ISBN = "979-8-89176-195-7"
} | |||||
| LLM Benchmarks | 134,375 question-option pairs | CC BY-SA 4.0 | 2025 | ↗✓ 2026-07-19 | |
Bengali massive multitask language understanding benchmark spanning 23 domains. Source: BnMMLU (2025). @inproceedings{joy-shatabda-2026-bnmmlu,
title = "{B}n{MMLU}: Measuring Massive Multitask Language Understanding in {B}engali",
author = "Joy, Saman Sarker and
Shatabda, Swakkhar",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.findings-acl.593/",
doi = "10.18653/v1/2026.findings-acl.593",
pages = "12211--12230",
ISBN = "979-8-89176-395-1"
} | |||||
| LLM Benchmarks | 8,792 problems | CC BY 4.0 | 2025 | ↗✓ 2026-07-26 | |
Bengali math word problems with step-by-step solutions adapted from the GSM8K train and test sets. Source: SOMADHAN (2025). @misc{paul2025leveraginglargelanguagemodels,
title={Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning},
author={Bidyarthi Paul and Jalisha Jashim Era and Mirazur Rahman Zim and Tahmid Sattar Aothoi and Faisal Muhammad Shah},
year={2025},
eprint={2505.21354},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.21354},
} | |||||
| Machine Translation | 42,705 pairs | MIT | 2024 | ↗✓ 2026-07-19 | |
Benchmark for back-transliteration of romanized Bangla to Bangla script. Source: BanglaTLit (Findings EMNLP 2024). @inproceedings{fahim-etal-2024-banglatlit,
title = "{B}angla{TL}it: A Benchmark Dataset for Back-Transliteration of {R}omanized {B}angla",
author = "Fahim, Md and
Shifat, Fariha Tanjim and
Haider, Fabiha and
Barua, Deeparghya Dutta and
Sourove, MD Sakib Ul Rahman and
Ishmam, Md Farhan and
Bhuiyan, Md Farhad Alam",
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2024",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.findings-emnlp.859/",
doi = "10.18653/v1/2024.findings-emnlp.859",
pages = "14656--14672"
} | |||||
| LLM Benchmarks | 5,161 questions | CC BY-NC-SA 4.0 | 2024 | ↗✓ 2026-07-19 | |
Parallel Bangla–English science exam questions for LLM evaluation. Source: BEnQA (2024). @article{shafayat2024benqa,
title = {{BE}n{QA}: A Question Answering and Reasoning Benchmark for {B}engali and {E}nglish},
author = {Shafayat, Sheikh and others},
journal = {arXiv preprint arXiv:2403.10900},
year = {2024}
} | |||||
| Sentiment Analysis | 20,000 samples | MIT | 2024 | ↗✓ 2026-07-19 | |
Bengali-English code-mixed sentiment dataset with 4 labels from social media and e-commerce. Source: BnSentMix (LoResLM 2025). @inproceedings{alam-etal-2025-bnsentmix,
title = "{B}n{S}ent{M}ix: A Diverse {B}engali-{E}nglish Code-Mixed Dataset for Sentiment Analysis",
author = "Alam, Sadia and
Ishmam, Md Farhan and
Alvee, Navid Hasin and
Siddique, Md Shahnewaz and
Hossain, Md Azam and
Kamal, Abu Raihan Mostofa",
editor = "Hettiarachchi, Hansi and
Ranasinghe, Tharindu and
Rayson, Paul and
Mitkov, Ruslan and
Gaber, Mohamed and
Premasiri, Damith and
Tan, Fiona Anting and
Uyangodage, Lasitha",
booktitle = "Proceedings of the First Workshop on Language Models for Low-Resource Languages",
month = jan,
year = "2025",
address = "Abu Dhabi, United Arab Emirates",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.loreslm-1.4/",
pages = "68--77"
} | |||||
| ASR / TTS | 1,200+ hrs recorded | CC0 | 2024 | ↗✓ 2026-07-01 | |
Continuously growing crowdsourced speech corpus. Source: Mozilla Common Voice. @inproceedings{ardila-etal-2020-common,
title = "Common Voice: A Massively-Multilingual Speech Corpus",
author = "Ardila, Rosana and
Branson, Megan and
Davis, Kelly and
Kohler, Michael and
Meyer, Josh and
Henretty, Michael and
Morais, Reuben and
Saunders, Lindsay and
Tyers, Francis and
Weber, Gregor",
editor = "Calzolari, Nicoletta and
B{\'e}chet, Fr{\'e}d{\'e}ric and
Blache, Philippe and
Choukri, Khalid and
Cieri, Christopher and
Declerck, Thierry and
Goggi, Sara and
Isahara, Hitoshi and
Maegaard, Bente and
Mariani, Joseph and
Mazo, H{\'e}l{\`e}ne and
Moreno, Asuncion and
Odijk, Jan and
Piperidis, Stelios",
booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference",
month = may,
year = "2020",
address = "Marseille, France",
publisher = "European Language Resources Association",
url = "https://aclanthology.org/2020.lrec-1.520/",
pages = "4218--4222",
language = "eng",
ISBN = "979-10-95546-34-4"
} | |||||
| Sentiment Analysis | 1,071 instances | GPL-3.0 | 2024 | ↗✓ 2026-07-19 | |
Bangla-English-Hindi code-mixed emotion detection test set. Source: EmoMix-3L (WILDRE 2024). @inproceedings{raihan-etal-2024-emomix,
title = "{E}mo{M}ix-3{L}: A Code-Mixed Dataset for {B}angla-{E}nglish-{H}indi for Emotion Detection",
author = "Raihan, Nishat and
Goswami, Dhiman and
Mahmud, Antara and
Anastasopoulos, Antonios and
Zampieri, Marcos",
editor = "Jha, Girish Nath and
L., Sobha and
Bali, Kalika and
Ojha, Atul Kr.",
booktitle = "Proceedings of the 7th Workshop on Indian Language Data: Resources and Evaluation",
month = may,
year = "2024",
address = "Torino, Italia",
publisher = "ELRA and ICCL",
url = "https://aclanthology.org/2024.wildre-1.2/",
pages = "11--16"
} | |||||
| LLM Benchmarks | 14,327 questions | Apache 2.0 | 2024 | ↗✓ 2026-07-19 | |
Culturally adapted MMLU; Bengali split. Source: Global-MMLU (2024). @misc{singh2025globalmmluunderstandingaddressing,
title={Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation},
author={Shivalika Singh and Angelika Romanou and Clémentine Fourrier and David I. Adelani and Jian Gang Ngui and Daniel Vila-Suero and Peerat Limkonchotiwat and Kelly Marchisio and Wei Qi Leong and Yosephine Susanto and Raymond Ng and Shayne Longpre and Wei-Yin Ko and Sebastian Ruder and Madeline Smith and Antoine Bosselut and Alice Oh and Andre F. T. Martins and Leshem Choshen and Daphne Ippolito and Enzo Ferrante and Marzieh Fadaee and Beyza Ermis and Sara Hooker},
year={2025},
eprint={2412.03304},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.03304},
} | |||||
| LLM Benchmarks | 22 datasets | Mixed | 2024 | ↗✓ 2026-08-23 | |
Multilingual LLM evaluation suite including Bangla tasks. Source: MEGAVERSE (NAACL 2024). @inproceedings{ahuja-etal-2024-megaverse,
title = "{MEGAVERSE}: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks",
author = "Ahuja, Sanchit and
Aggarwal, Divyanshu and
Gumma, Varun and
Watts, Ishaan and
Sathe, Ashutosh and
Ochieng, Millicent and
Hada, Rishav and
Jain, Prachi and
Ahmed, Mohamed and
Bali, Kalika and
Sitaram, Sunayana",
editor = "Duh, Kevin and
Gomez, Helena and
Bethard, Steven",
booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = jun,
year = "2024",
address = "Mexico City, Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.naacl-long.143/",
doi = "10.18653/v1/2024.naacl-long.143",
pages = "2598--2637"
} | |||||
| Sentiment Analysis | 7,058 instances | MIT | 2024 | ↗✓ 2026-07-19 | |
Bengali political sentiment dataset compiled from election-period news, labelled positive/negative. Source: Motamot (2024). @inproceedings{Johora_Faria_2024, title={Motamot: A Dataset for Revealing the Supremacy of Large Language Models Over Transformer Models in Bengali Political Sentiment Analysis}, url={http://dx.doi.org/10.1109/TENSYMP61132.2024.10752197}, DOI={10.1109/tensymp61132.2024.10752197}, booktitle={2024 IEEE Region 10 Symposium (TENSYMP)}, publisher={IEEE}, author={Johora Faria, Fatema Tuj and Moin, Mukaffi Bin and Mumu, Rabeya Islam and Alam Abir, Md Mahabubul and Alfy, Abrar Nawar and Alam, Mohammad Shafiul}, year={2024}, month=Sept, pages={1–8} } | |||||
| Named Entity Recognition | 22,144 sentences | MIT | 2023 | ↗✓ 2026-07-19 | |
Largest gold-standard Bangla NER corpus, 8 entity types. Source: B-NER (IEEE Access 2023). @article{Haque_2023, title={B-NER: A Novel Bangla Named Entity Recognition Dataset With Largest Entities and Its Baseline Evaluation}, volume={11}, ISSN={2169-3536}, url={http://dx.doi.org/10.1109/ACCESS.2023.3267746}, DOI={10.1109/access.2023.3267746}, journal={IEEE Access}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Haque, Md. Zahidul and Zaman, Sakib and Saurav, Jillur Rahman and Haque, Summit and Islam, Md. Saiful and Amin, Mohammad Ruhul}, year={2023}, pages={45194–45205} } | |||||
| Sentiment Analysis | 158,065 reviews | CC BY-NC-SA 4.0 | 2023 | ↗✓ 2026-07-19 | |
Large-scale Bangla book-review sentiment dataset in three classes (positive, negative, neutral). Source: BanglaBook (Findings ACL 2023). @inproceedings{kabir-etal-2023-banglabook,
title = "{B}angla{B}ook: A Large-scale {B}angla Dataset for Sentiment Analysis from Book Reviews",
author = "Kabir, Mohsinul and
Bin Mahfuz, Obayed and
Raiyan, Syed Rifat and
Mahmud, Hasan and
Hasan, Md Kamrul",
editor = "Rogers, Anna and
Boyd-Graber, Jordan and
Okazaki, Naoaki",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.findings-acl.80/",
doi = "10.18653/v1/2023.findings-acl.80",
pages = "1237--1247"
} | |||||
| Summarization | 2,350 pairs | CC BY-NC-SA 4.0 | 2023 | ↗✓ 2026-07-19 | |
Abstractive summarization dataset for Bangla consumer health questions. Source: BanglaCHQ-Summ (BLP 2023). @inproceedings{khan-etal-2023-banglachq,
title = "{B}angla{CHQ}-Summ: An Abstractive Summarization Dataset for Medical Queries in {B}angla Conversational Speech",
author = "Khan, Alvi and
Kamal, Fida and
Chowdhury, Mohammad Abrar and
Ahmed, Tasnim and
Laskar, Md Tahmid Rahman and
Ahmed, Sabbir",
editor = "Alam, Firoj and
Kar, Sudipta and
Chowdhury, Shammur Absar and
Sadeque, Farig and
Amin, Ruhul",
booktitle = "Proceedings of the First Workshop on Bangla Language Processing (BLP-2023)",
month = dec,
year = "2023",
address = "Singapore",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.banglalp-1.10/",
doi = "10.18653/v1/2023.banglalp-1.10",
pages = "85--93"
} | |||||
| Speech Emotion Recognition | 792 utterances | CC BY 4.0 | 2023 | ↗✓ 2026-07-22 | |
Bangla emotional speech; 22 speakers (11M/11F), 6 emotions. Source: Sultana et al. / Mendeley Data. @article{Sultana_2025, title={BanSpEmo: a Bangla audio dataset for speech emotion recognition and its baseline evaluation}, volume={37}, ISSN={2502-4752}, url={http://dx.doi.org/10.11591/ijeecs.v37.i3.pp2044-2057}, DOI={10.11591/ijeecs.v37.i3.pp2044-2057}, number={3}, journal={Indonesian Journal of Electrical Engineering and Computer Science}, publisher={Institute of Advanced Engineering and Science}, author={Sultana, Babe and Hussain, Md Gulzar and Rahman, Mahmuda}, year={2025}, month=Mar, pages={2044} } | |||||
| Speech Emotion Recognition | 900 recordings | CC BY 4.0 | 2023 | ↗✓ 2026-07-26 | |
Realistic Bangla emotional speech from 35 actors, with low/high intensity labels for 4 of 5 emotions. Source: KBES (Data in Brief 2023). @article{Billah_2023, title={KBES: A dataset for realistic Bangla speech emotion recognition with intensity level}, volume={51}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2023.109741}, DOI={10.1016/j.dib.2023.109741}, journal={Data in Brief}, publisher={Elsevier BV}, author={Billah, Md. Masum and Sarker, Md. Likhon and Akhand, M. A. H.}, year={2023}, month=Dec, pages={109741} } | |||||
| Hate Speech Detection | 1,001 instances | AGPL-3.0 | 2023 | ↗✓ 2026-07-19 | |
Bangla-English-Hindi code-mixed offensive language identification test set. Source: OffMix-3L (SocialNLP 2023). @inproceedings{goswami-etal-2023-offmix,
title = "{O}ff{M}ix-3{L}: A Novel Code-Mixed Test Dataset in {B}angla-{E}nglish-{H}indi for Offensive Language Identification",
author = "Goswami, Dhiman and
Raihan, Md Nishat and
Mahmud, Antara and
Anastasopoulos, Antonios and
Zampieri, Marcos",
editor = "Ku, Lun-Wei and
Li, Cheng-Te",
booktitle = "Proceedings of the 11th International Workshop on Natural Language Processing for Social Media",
month = nov,
year = "2023",
address = "Bali, Indonesia",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.socialnlp-1.3/",
doi = "10.18653/v1/2023.socialnlp-1.3",
pages = "21--27"
} | |||||
| ASR / TTS | 1,177 hrs | CC BY-NC 4.0 | 2023 | ↗✓ 2026-07-19 | |
Large-scale out-of-distribution Bangla ASR benchmark. Source: OOD-Speech (2023). @misc{rakib2023oodspeechlargebengalispeech,
title={OOD-Speech: A Large Bengali Speech Recognition Dataset for Out-of-Distribution Benchmarking},
author={Fazle Rabbi Rakib and Souhardya Saha Dip and Samiul Alam and Nazia Tasnim and Md. Istiak Hossain Shihab and Md. Nazmuddoha Ansary and Syed Mobassir Hossen and Marsia Haque Meghla and Mamunur Mamun and Farig Sadeque and Sayma Sultana Chowdhury and Tahsin Reasat and Asif Sushmit and Ahmed Imtiaz Humayun},
year={2023},
eprint={2305.09688},
archivePrefix={arXiv},
primaryClass={eess.AS},
url={https://arxiv.org/abs/2305.09688},
} | |||||
| Sentiment Analysis | 1,007 instances | AGPL-3.0 | 2023 | ↗✓ 2026-07-19 | |
Gold-standard Bangla-English-Hindi code-mixed sentiment analysis test set. Source: SentMix-3L (SEALP 2023). @inproceedings{raihan-etal-2023-sentmix,
title = "{S}ent{M}ix-3{L}: A Novel Code-Mixed Test Dataset in {B}angla-{E}nglish-{H}indi for Sentiment Analysis",
author = "Raihan, Md Nishat and
Goswami, Dhiman and
Mahmud, Antara and
Anastasopoulos, Antonios and
Zampieri, Marcos",
editor = "Wijaya, Derry and
Aji, Alham Fikri and
Vania, Clara and
Winata, Genta Indra and
Purwarianti, Ayu",
booktitle = "Proceedings of the First Workshop in South East Asian Language Processing",
month = nov,
year = "2023",
address = "Nusa Dua, Bali, Indonesia",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.sealp-1.6/",
doi = "10.18653/v1/2023.sealp-1.6",
pages = "79--84"
} | |||||
| Hate Speech Detection | 5,000 comments | CC BY 4.0 | 2023 | ↗✓ 2026-03-08 | |
Offensive language in transliterated (Banglish) text. Source: Transliterated offensive language ID. @inproceedings{raihan-etal-2023-offensive,
title = "Offensive Language Identification in Transliterated and Code-Mixed {B}angla",
author = "Raihan, Md Nishat and
Tanmoy, Umma and
Islam, Anika Binte and
North, Kai and
Ranasinghe, Tharindu and
Anastasopoulos, Antonios and
Zampieri, Marcos",
editor = "Alam, Firoj and
Kar, Sudipta and
Chowdhury, Shammur Absar and
Sadeque, Farig and
Amin, Ruhul",
booktitle = "Proceedings of the First Workshop on Bangla Language Processing (BLP-2023)",
month = dec,
year = "2023",
address = "Singapore",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.banglalp-1.1/",
doi = "10.18653/v1/2023.banglalp-1.1",
pages = "1--6"
} | |||||
| Machine Translation | 32,500 sentences | CC BY 4.0 | 2023 | ↗✓ 2026-07-19 | |
Translation benchmark from five Bangla regional dialects to standard Bangla. Source: Vashantor (2023). @misc{faria2025vashantorlargescalemultilingualbenchmark,
title={Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language},
author={Fatema Tuj Johora Faria and Mukaffi Bin Moin and Ahmed Al Wase and Mehidi Ahmmed and Md. Rabius Sani and Tashreef Muhammad},
year={2025},
eprint={2311.11142},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2311.11142},
} | |||||
| Machine Translation | 466,630 pairs | CC BY-NC-SA 4.0 | 2022 | ↗✓ 2026-07-19 | |
High-quality synthetic paraphrase pairs. Source: BanglaParaphrase (AACL 2022). @inproceedings{akil-etal-2022-banglaparaphrase,
title = "{B}angla{P}araphrase: A High-Quality {B}angla Paraphrase Dataset",
author = "Akil, Ajwad and
Sultana, Najrin and
Bhattacharjee, Abhik and
Shahriyar, Rifat",
editor = "He, Yulan and
Ji, Heng and
Li, Sujian and
Liu, Yang and
Chang, Chua-Hui",
booktitle = "Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)",
month = nov,
year = "2022",
address = "Online only",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.aacl-short.33/",
doi = "10.18653/v1/2022.aacl-short.33",
pages = "261--272"
} | |||||
| Question Answering | 14,889 QA pairs | CC BY-NC 4.0 | 2022 | ↗✓ 2026-06-15 | |
Human-annotated reading comprehension with answerable/unanswerable questions. Source: BanglaRQA (Findings EMNLP 2022). @inproceedings{ekram-etal-2022-banglarqa,
title = "{B}angla{RQA}: A Benchmark Dataset for Under-resourced {B}angla Language Reading Comprehension-based Question Answering with Diverse Question-Answer Types",
author = "Ekram, Syed Mohammed Sartaj and
Rahman, Adham Arik and
Altaf, Md. Sajid and
Islam, Mohammed Saidul and
Rahman, Mehrab Mustafy and
Rahman, Md Mezbaur and
Hossain, Md Azam and
Kamal, Abu Raihan Mostofa",
editor = "Goldberg, Yoav and
Kozareva, Zornitsa and
Zhang, Yue",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2022",
month = dec,
year = "2022",
address = "Abu Dhabi, United Arab Emirates",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-emnlp.186/",
doi = "10.18653/v1/2022.findings-emnlp.186",
pages = "2518--2532"
} | |||||
| Speech Emotion Recognition | 1,467 recordings | CC BY 4.0 | 2022 | ↗✓ 2026-07-26 | |
Realistic Bangla emotional speech from 34 speakers, balanced by gender, across 5 emotions. Source: BanglaSER (Data in Brief 2022). @article{Das_2022, title={BanglaSER: A speech emotion recognition dataset for the Bangla language}, volume={42}, ISSN={2352-3409}, url={http://dx.doi.org/10.1016/j.dib.2022.108091}, DOI={10.1016/j.dib.2022.108091}, journal={Data in Brief}, publisher={Elsevier BV}, author={Das, Rakesh Kumar and Islam, Nahidul and Ahmed, Md. Rayhan and Islam, Salekul and Shatabda, Swakkhar and Islam, A.K.M. Muzahidul}, year={2022}, month=June, pages={108091} } | |||||
| Hate Speech Detection | 50,281 comments | CC BY-NC 4.0 | 2022 | ↗✓ 2026-06-05 | |
Hierarchically annotated hate speech: target, type, severity. Source: BD-SHS (LREC 2022). @inproceedings{romim-etal-2022-bd,
title = "{BD}-{SHS}: A Benchmark Dataset for Learning to Detect Online {B}angla Hate Speech in Different Social Contexts",
author = "Romim, Nauros and
Ahmed, Mosahed and
Islam, Md Saiful and
Sen Sharma, Arnab and
Talukder, Hriteshwar and
Amin, Mohammad Ruhul",
editor = "Calzolari, Nicoletta and
B{\'e}chet, Fr{\'e}d{\'e}ric and
Blache, Philippe and
Choukri, Khalid and
Cieri, Christopher and
Declerck, Thierry and
Goggi, Sara and
Isahara, Hitoshi and
Maegaard, Bente and
Mariani, Joseph and
Mazo, H{\'e}l{\`e}ne and
Odijk, Jan and
Piperidis, Stelios",
booktitle = "Proceedings of the Thirteenth Language Resources and Evaluation Conference",
month = jun,
year = "2022",
address = "Marseille, France",
publisher = "European Language Resources Association",
url = "https://aclanthology.org/2022.lrec-1.552/",
pages = "5153--5162"
} | |||||
| Sentiment Analysis | 7,000 texts | Research only | 2022 | ↗✓ 2026-07-19 | |
Six-class emotion corpus (anger, disgust, fear, joy, sadness, surprise). Source: Emotion classification corpus. @Article{Iqbal2022,
author="Iqbal, MD. Asif
and Das, Avishek
and Sharif, Omar
and Hoque, Mohammed Moshiul
and Sarker, Iqbal H.",
title="BEmoC: A Corpus for Identifying Emotion in Bengali Texts",
journal="SN Computer Science",
year="2022",
month="Jan",
day="17",
volume="3",
number="2",
pages="135",
abstract="Emotion classification in text has growing interest among NLP experts due to the enormous availability of people's emotions and its emergence on various Web 2.0 applications/services. Emotion classification in the Bengali texts is also gradually being considered as an important task for sports, e-commerce, entertainments, and security applications. However, It is a very critical task to develop an automatic emotion classification system for low-resource languages such as, Bengali. Scarcity of resources and deficiency of benchmark corpora make the task more complicated. Thus, the development of a benchmark corpus is the prerequisite to develop an emotion classifier for Bengali texts. This paper describes the development of an emotional corpus (hereafter called `BEmoC') for classifying six emotions in Bengali texts. The corpus development process consists of four key steps: data crawling, pre-processing, labelling, and verification. A total of 7000 texts are labelled into six basic emotion categories such as anger, fear, surprise, sadness, joy, and disgust, respectively. Dataset evaluation with 0.969 Cohen's $\kappa$ score indicates the close agreement between the corpus annotators and the expert. The analysis of evaluation also represents that the distribution of emotion words obeys Zipf's law. Moreover, the results of BEmoC analysis shown in terms of coding reliability, emotion density, and most frequent emotion words, respectively.",
issn="2661-8907",
doi="10.1007/s42979-022-01028-w",
url="https://doi.org/10.1007/s42979-022-01028-w"
} | |||||
| Machine Translation | 2,009 public sentences | CC BY-SA 4.0 | 2022 | ↗✓ 2026-09-08 | |
Bengali ben_Beng evaluation data with 997 dev and 1,012 devtest sentences; excludes the hidden test split. Source: No Language Left Behind (2022). @misc{nllbteam2022languageleftbehindscaling,
title={No Language Left Behind: Scaling Human-Centered Machine Translation},
author={NLLB Team and Marta R. Costa-jussà and James Cross and Onur Çelebi and Maha Elbayad and Kenneth Heafield and Kevin Heffernan and Elahe Kalbassi and Janice Lam and Daniel Licht and Jean Maillard and Anna Sun and Skyler Wang and Guillaume Wenzek and Al Youngblood and Bapi Akula and Loic Barrault and Gabriel Mejia Gonzalez and Prangthip Hansanti and John Hoffman and Semarley Jarrett and Kaushik Ram Sadagopan and Dirk Rowe and Shannon Spruit and Chau Tran and Pierre Andrews and Necip Fazil Ayan and Shruti Bhosale and Sergey Edunov and Angela Fan and Cynthia Gao and Vedanuj Goswami and Francisco Guzmán and Philipp Koehn and Alexandre Mourachko and Christophe Ropers and Safiyyah Saleem and Holger Schwenk and Jeff Wang},
year={2022},
eprint={2207.04672},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2207.04672},
} | |||||
| Named Entity Recognition | 15,300 sentences | CC BY 4.0 | 2022 | ↗✓ 2026-06-10 | |
Complex NER: creative works, groups, products; Bangla track. Source: SemEval-2022 Task 11. @inproceedings{malmasi-etal-2022-semeval,
title = "{S}em{E}val-2022 Task 11: Multilingual Complex Named Entity Recognition ({M}ulti{C}o{NER})",
author = "Malmasi, Shervin and
Fang, Anjie and
Fetahu, Besnik and
Kar, Sudipta and
Rokhlenko, Oleg",
editor = "Emerson, Guy and
Schluter, Natalie and
Stanovsky, Gabriel and
Kumar, Ritesh and
Palmer, Alexis and
Schneider, Nathan and
Singh, Siddharth and
Ratan, Shyam",
booktitle = "Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022)",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.semeval-1.196/",
doi = "10.18653/v1/2022.semeval-1.196",
pages = "1412--1437"
} | |||||
| Text Classification | 665k articles | CC BY 4.0 | 2022 | ↗✓ 2026-01-30 | |
Large single-label news corpus, 8 categories, 6 dailies. Source: Potrika (2022). @misc{ahmad2022potrikarawbalancednewspaper,
title={Potrika: Raw and Balanced Newspaper Datasets in the Bangla Language with Eight Topics and Five Attributes},
author={Istiak Ahmad and Fahad AlQurashi and Rashid Mehmood},
year={2022},
eprint={2210.09389},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2210.09389},
} | |||||
| Machine Translation | 8.6M pairs | CC0 | 2022 | ↗✓ 2026-07-19 | |
Largest publicly available bn–en parallel collection. Source: Samanantar (TACL 2022). @article{ramesh-etal-2022-samanantar,
title = "Samanantar: The Largest Publicly Available Parallel Corpora Collection for 11 {I}ndic Languages",
author = "Ramesh, Gowtham and
Doddapaneni, Sumanth and
Bheemaraj, Aravinth and
Jobanputra, Mayank and
AK, Raghavan and
Sharma, Ajitesh and
Sahoo, Sujit and
Diddee, Harshita and
J, Mahalakshmi and
Kakwani, Divyanshu and
Kumar, Navneet and
Pradeep, Aswin and
Nagaraj, Srihari and
Deepak, Kumar and
Raghavan, Vivek and
Kunchukuttan, Anoop and
Kumar, Pratyush and
Khapra, Mitesh Shantadevi",
editor = "Roark, Brian and
Nenkova, Ani",
journal = "Transactions of the Association for Computational Linguistics",
volume = "10",
year = "2022",
address = "Cambridge, MA",
publisher = "MIT Press",
url = "https://aclanthology.org/2022.tacl-1.9/",
doi = "10.1162/tacl_a_00452",
pages = "145--162"
} | |||||
| Summarization | 19,096 pairs | Research only | 2021 | ↗✓ 2026-02-20 | |
News summarization pairs crawled from Bangla dailies. Source: Bangla abstractive news summarization. @InProceedings{10.1007/978-981-33-4673-4_4,
author="Bhattacharjee, Prithwiraj
and Mallick, Avi
and Saiful Islam, Md.
and Marium-E-Jannat",
editor="Kaiser, M. Shamim
and Bandyopadhyay, Anirban
and Mahmud, Mufti
and Ray, Kanad",
title="Bengali Abstractive News Summarization (BANS): A Neural Attention Approach",
booktitle="Proceedings of International Conference on Trends in Computational and Cognitive Engineering",
year="2021",
publisher="Springer Singapore",
address="Singapore",
pages="41--51",
abstract="Bhattacharjee, PrithwirajMallick, AviSaiful Islam, Md.Marium-E-JannatAbstractive summarization is the process of generating novel sentences based on the information extracted from the original text document while retaining the context. Due to abstractive summarization's underlying complexities, most of the past research work has been done on the extractive summarization approach. Nevertheless, with the triumph of the sequence-to-sequence (seq2seq) model, abstractive summarization becomes more viable. Although a significant number of notable research has been done in the English language based on abstractive summarization, only a couple of works have been done on Bengali abstractive news summarization (BANS). In this article, we presented a seq2seq based Long Short-Term Memory (LSTM) network model with attention at encoder-decoder. Our proposed system deploys a local attention-based model that produces a long sequence of words with lucid and human-like generated sentences with noteworthy information of the original document. We also prepared a dataset of more than 19 k articles and corresponding human-written summaries collected from bangla.bdnews24.com (https://bangla.bdnews24.com/) which is till now the most extensive dataset for Bengali news document summarization and publicly published in Kaggle (https://www.kaggle.com/prithwirajsust/bengali-news-summarization-dataset) We evaluated our model qualitatively and quantitatively and compared it with other published results. It showed significant improvement in terms of human evaluation scores with state-of-the-art approaches for BANS.",
isbn="978-981-33-4673-4"
} | |||||
| Sentiment Analysis | 15,728 comments | CC BY-ND 4.0 | 2021 | ↗✓ 2026-09-08 | |
Sentiment on noisy user comments from news and social platforms across 13 domains. Source: SentNoB (Findings EMNLP 2021). @inproceedings{islam-etal-2021-sentnob-dataset,
title = "{S}ent{N}o{B}: A Dataset for Analysing Sentiment on Noisy {B}angla Texts",
author = "Islam, Khondoker Ittehadul and
Kar, Sudipta and
Islam, Md Saiful and
Amin, Mohammad Ruhul",
editor = "Moens, Marie-Francine and
Huang, Xuanjing and
Specia, Lucia and
Yih, Scott Wen-tau",
booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
month = nov,
year = "2021",
address = "Punta Cana, Dominican Republic",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2021.findings-emnlp.278/",
doi = "10.18653/v1/2021.findings-emnlp.278",
pages = "3265--3271"
} | |||||
| Question Answering | 132,777 QA pairs | CC BY-NC-SA 4.0 | 2021 | ↗✓ 2026-07-19 | |
Machine-translated + human-curated SQuAD 2.0 for Bangla. Source: BanglaBERT benchmarks. @inproceedings{bhattacharjee-etal-2022-banglabert,
title = "{B}angla{BERT}: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in {B}angla",
author = "Bhattacharjee, Abhik and
Hasan, Tahmid and
Ahmad, Wasi and
Mubasshir, Kazi Samin and
Islam, Md Saiful and
Iqbal, Anindya and
Rahman, M. Sohel and
Shahriyar, Rifat",
editor = "Carpuat, Marine and
de Marneffe, Marie-Catherine and
Meza Ruiz, Ivan Vladimir",
booktitle = "Findings of the Association for Computational Linguistics: NAACL 2022",
month = jul,
year = "2022",
address = "Seattle, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.findings-naacl.98/",
doi = "10.18653/v1/2022.findings-naacl.98",
pages = "1318--1327"
} | |||||
| Speech Emotion Recognition | 7,000 utterances / 7.7 hrs | CC BY 4.0 | 2021 | ↗✓ 2026-07-22 | |
Audio-only Bangla emotional speech; 20 actors (gender-balanced), 7 emotions. Source: Sultana et al. (SUST) / Zenodo. @article{10.1371/journal.pone.0250173,
doi = {10.1371/journal.pone.0250173},
author = {Sultana, Sadia AND Rahman, M. Shahidur AND Selim, M. Reza AND Iqbal, M. Zafar},
journal = {PLOS ONE},
publisher = {Public Library of Science},
title = {SUST Bangla Emotional Speech Corpus (SUBESCO): An audio-only emotional speech corpus for Bangla},
year = {2021},
month = {04},
volume = {16},
url = {https://doi.org/10.1371/journal.pone.0250173},
pages = {1-27},
abstract = {SUBESCO is an audio-only emotional speech corpus for Bangla language. The total duration of the corpus is in excess of 7 hours containing 7000 utterances, and it is the largest emotional speech corpus available for this language. Twenty native speakers participated in the gender-balanced set, each recording of 10 sentences simulating seven targeted emotions. Fifty university students participated in the evaluation of this corpus. Each audio clip of this corpus, except those of Disgust emotion, was validated four times by male and female raters. Raw hit rates and unbiased rates were calculated producing scores above chance level of responses. Overall recognition rate was reported to be above 70% for human perception tests. Kappa statistics and intra-class correlation coefficient scores indicated high-level of inter-rater reliability and consistency of this corpus evaluation. SUBESCO is an Open Access database, licensed under Creative Common Attribution 4.0 International, and can be downloaded free of charge from the web link: https://doi.org/10.5281/zenodo.4526477.},
number = {4},
} | |||||
| POS Tagging | 320 tokens / 56 sentences | CC BY-SA 4.0 | 2021 | ↗✓ 2026-08-23 | |
Tiny Universal Dependencies sample treebank with UPOS and dependency relations. Citation covers the UD 2.9 release, including this treebank. Source: Universal Dependencies (since v2.9). @misc{11234/1-4611,
title = {Universal Dependencies 2.9},
author = {Zeman, Daniel and Nivre, Joakim and Abrams, Mitchell and Ackermann, Elia and Aepli, No{\"e}mi and Aghaei, Hamid and Agi{\'c}, {\v Z}eljko and Ahmadi, Amir and Ahrenberg, Lars and Ajede, Chika Kennedy and Aleksandravi{\v c}i{\=u}t{\.e}, Gabriel{\.e} and Alfina, Ika and Antonsen, Lene and Aplonova, Katya and Aquino, Angelina and Aragon, Carolina and Aranzabe, Maria Jesus and Ar{\i}can, Bilge Nas and Arnard{\'o}ttir, {\t H}{\'o}runn and Arutie, Gashaw and Arwidarasti, Jessica Naraiswari and Asahara, Masayuki and Aslan, Deniz Baran and Ateyah, Luma and Atmaca, Furkan and Attia, Mohammed and Atutxa, Aitziber and Augustinus, Liesbeth and Badmaeva, Elena and Balasubramani, Keerthana and Ballesteros, Miguel and Banerjee, Esha and Bank, Sebastian and Barbu Mititelu, Verginica and Barkarson, Starkaður and Basile, Rodolfo and Basmov, Victoria and Batchelor, Colin and Bauer, John and Bedir, Seyyit Talha and Bengoetxea, Kepa and Berk, G{\"o}zde and Berzak, Yevgeni and Bhat, Irshad Ahmad and Bhat, Riyaz Ahmad and Biagetti, Erica and Bick, Eckhard and Bielinskien{\.e}, Agn{\.e} and Bjarnad{\'o}ttir, Krist{\'{\i}}n and Blokland, Rogier and Bobicev, Victoria and Boizou, Lo{\"{\i}}c and Borges V{\"o}lker, Emanuel and B{\"o}rstell, Carl and Bosco, Cristina and Bouma, Gosse and Bowman, Sam and Boyd, Adriane and Braggaar, Anouck and Brokait{\.e}, Kristina and Burchardt, Aljoscha and Candito, Marie and Caron, Bernard and Caron, Gauthier and Cassidy, Lauren and Cavalcanti, Tatiana and Cebiro{\u g}lu Eryi{\u g}it, G{\"u}l{\c s}en and Cecchini, Flavio Massimiliano and Celano, Giuseppe G. A. and {\v C}{\'e}pl{\"o}, Slavom{\'{\i}}r and Cesur, Neslihan and Cetin, Savas and {\c C}etino{\u g}lu, {\"O}zlem and Chalub, Fabricio and Chauhan, Shweta and Chi, Ethan and Chika, Taishi and Cho, Yongseok and Choi, Jinho and Chun, Jayeol and Chung, Juyeon and Cignarella, Alessandra T. and Cinkov{\'a}, Silvie and Collomb, Aur{\'e}lie and {\c C}{\"o}ltekin, {\c C}a{\u g}r{\i} and Connor, Miriam and Courtin, Marine and Cristescu, Mihaela and Daniel, Philemon and Davidson, Elizabeth and de Marneffe, Marie-Catherine and de Paiva, Valeria and Derin, Mehmet Oguz and de Souza, Elvis and Diaz de Ilarraza, Arantza and Dickerson, Carly and Dinakaramani, Arawinda and Di Nuovo, Elisa and Dione, Bamba and Dirix, Peter and Dobrovoljc, Kaja and Dozat, Timothy and Droganova, Kira and Dwivedi, Puneet and Eckhoff, Hanne and Eiche, Sandra and Eli, Marhaba and Elkahky, Ali and Ephrem, Binyam and Erina, Olga and Erjavec, Toma{\v z} and Etienne, Aline and Evelyn, Wograine and Facundes, Sidney and Farkas, Rich{\'a}rd and Ferdaousi, Jannatul and Fernanda, Mar{\'{\i}}lia and Fernandez Alcalde, Hector and Foster, Jennifer and Freitas, Cl{\'a}udia and Fujita, Kazunori and Gajdo{\v s}ov{\'a}, Katar{\'{\i}}na and Galbraith, Daniel and Garcia, Marcos and G{\"a}rdenfors, Moa and Garza, Sebastian and Gerardi, Fabr{\'{\i}}cio Ferraz and Gerdes, Kim and Ginter, Filip and Godoy, Gustavo and Goenaga, Iakes and Gojenola, Koldo and G{\"o}k{\i}rmak, Memduh and Goldberg, Yoav and G{\'o}mez Guinovart, Xavier and Gonz{\'a}lez Saavedra, Berta and Grici{\=u}t{\.e}, Bernadeta and Grioni, Matias and Grobol, Lo{\"{\i}}c and Gr{\=u}z{\={\i}}tis, Normunds and Guillaume, Bruno and Guillot-Barbance, C{\'e}line and G{\"u}ng{\"o}r, Tunga and Habash, Nizar and Hafsteinsson, Hinrik and Haji{\v c}, Jan and Haji{\v c} jr., Jan and H{\"a}m{\"a}l{\"a}inen, Mika and H{\`a} M{\~y}, Linh and Han, Na-Rae and Hanifmuti, Muhammad Yudistira and Hardwick, Sam and Harris, Kim and Haug, Dag and Heinecke, Johannes and Hellwig, Oliver and Hennig, Felix and Hladk{\'a}, Barbora and Hlav{\'a}{\v c}ov{\'a}, Jaroslava and Hociung, Florinel and Hohle, Petter and Huber, Eva and Hwang, Jena and Ikeda, Takumi and Ingason, Anton Karl and Ion, Radu and Irimia, Elena and Ishola, {\d O}l{\'a}j{\'{\i}}d{\'e} and Ito, Kaoru and Jannat, Siratun and Jel{\'{\i}}nek, Tom{\'a}{\v s} and Jha, Apoorva and Johannsen, Anders and J{\'o}nsd{\'o}ttir, Hildur and J{\o}rgensen, Fredrik and Juutinen, Markus and K, Sarveswaran and Ka{\c s}{\i}kara, H{\"u}ner and Kaasen, Andre and Kabaeva, Nadezhda and Kahane, Sylvain and Kanayama, Hiroshi and Kanerva, Jenna and Kara, Neslihan and Katz, Boris and Kayadelen, Tolga and Kenney, Jessica and Kettnerov{\'a}, V{\'a}clava and Kirchner, Jesse and Klementieva, Elena and Klyachko, Elena and K{\"o}hn, Arne and K{\"o}ksal, Abdullatif and Kopacewicz, Kamil and Korkiakangas, Timo and K{\"o}se, Mehmet and Kotsyba, Natalia and Kovalevskait{\.e}, Jolanta and Krek, Simon and Krishnamurthy, Parameswari and K{\"u}bler, Sandra and Kuyruk{\c c}u, O{\u g}uzhan and Kuzgun, Asl{\i} and Kwak, Sookyoung and Laippala, Veronika and Lam, Lucia and Lambertino, Lorenzo and Lando, Tatiana and Larasati, Septina Dian and Lavrentiev, Alexei and Lee, John and L{\^e} H{\`{\^o}}ng, Phương and Lenci, Alessandro and Lertpradit, Saran and Leung, Herman and Levina, Maria and Li, Cheuk Ying and Li, Josie and Li, Keying and Li, Yuan and Lim, {KyungTae} and Lima Padovani, Bruna and Lind{\'e}n, Krister and Ljube{\v s}i{\'c}, Nikola and Loginova, Olga and Lusito, Stefano and Luthfi, Andry and Luukko, Mikko and Lyashevskaya, Olga and Lynn, Teresa and Macketanz, Vivien and Mahamdi, Menel and Maillard, Jean and Makazhanov, Aibek and Mandl, Michael and Manning, Christopher and Manurung, Ruli and Mar{\c s}an, B{\"u}{\c s}ra and M{\u a}r{\u a}nduc, C{\u a}t{\u a}lina and Mare{\v c}ek, David and Marheinecke, Katrin and Mart{\'{\i}}nez Alonso, H{\'e}ctor and Mart{\'{\i}}n-Rodr{\'{\i}}guez, Lorena and Martins, Andr{\'e} and Ma{\v s}ek, Jan and Matsuda, Hiroshi and Matsumoto, Yuji and Mazzei, Alessandro and {McDonald}, Ryan and {McGuinness}, Sarah and Mendon{\c c}a, Gustavo and Merzhevich, Tatiana and Miekka, Niko and Mischenkova, Karina and Misirpashayeva, Margarita and Missil{\"a}, Anna and Mititelu, C{\u a}t{\u a}lin and Mitrofan, Maria and Miyao, Yusuke and Mojiri Foroushani, {AmirHossein} and Moln{\'a}r, Judit and Moloodi, Amirsaeid and Montemagni, Simonetta and More, Amir and Moreno Romero, Laura and Moretti, Giovanni and Mori, Keiko Sophie and Mori, Shinsuke and Morioka, Tomohiko and Moro, Shigeki and Mortensen, Bjartur and Moskalevskyi, Bohdan and Muischnek, Kadri and Munro, Robert and Murawaki, Yugo and M{\"u}{\"u}risep, Kaili and Nainwani, Pinkey and Nakhl{\'e}, Mariam and Navarro Hor{\~n}iacek, Juan Ignacio and Nedoluzhko, Anna and Ne{\v s}pore-B{\=e}rzkalne, Gunta and Nevaci, Manuela and Nguy{\~{\^e}}n Th{\d i}, Lương and Nguy{\~{\^e}}n Th{\d i} Minh, Huy{\`{\^e}}n and Nikaido, Yoshihiro and Nikolaev, Vitaly and Nitisaroj, Rattima and Nourian, Alireza and Nurmi, Hanna and Ojala, Stina and Ojha, Atul Kr. and Ol{\'u}{\`o}kun, Ad{\'e}day{\d o}̀ and Omura, Mai and Onwuegbuzia, Emeka and Osenova, Petya and {\"O}stling, Robert and {\O}vrelid, Lilja and {\"O}zate{\c s}, {\c S}aziye Bet{\"u}l and {\"O}z{\c c}elik, Merve and {\"O}zg{\"u}r, Arzucan and {\"O}zt{\"u}rk Ba{\c s}aran, Balk{\i}z and Park, Hyunji Hayley and Partanen, Niko and Pascual, Elena and Passarotti, Marco and Patejuk, Agnieszka and Paulino-Passos, Guilherme and Peljak-{\L}api{\'n}ska, Angelika and Peng, Siyao and Perez, Cenel-Augusto and Perkova, Natalia and Perrier, Guy and Petrov, Slav and Petrova, Daria and Phelan, Jason and Piitulainen, Jussi and Pirinen, Tommi A and Pitler, Emily and Plank, Barbara and Poibeau, Thierry and Ponomareva, Larisa and Popel, Martin and Pretkalni{\c n}a, Lauma and Pr{\'e}vost, Sophie and Prokopidis, Prokopis and Przepi{\'o}rkowski, Adam and Puolakainen, Tiina and Pyysalo, Sampo and Qi, Peng and R{\"a}{\"a}bis, Andriela and Rademaker, Alexandre and Rahoman, Mizanur and Rama, Taraka and Ramasamy, Loganathan and Ramisch, Carlos and Rashel, Fam and Rasooli, Mohammad Sadegh and Ravishankar, Vinit and Real, Livy and Rebeja, Petru and Reddy, Siva and Regnault, Mathilde and Rehm, Georg and Riabov, Ivan and Rie{\ss}ler, Michael and Rimkut{\.e}, Erika and Rinaldi, Larissa and Rituma, Laura and Rizqiyah, Putri and Rocha, Luisa and R{\"o}gnvaldsson, Eir{\'{\i}}kur and Romanenko, Mykhailo and Rosa, Rudolf and Roșca, Valentin and Rovati, Davide and Rudina, Olga and Rueter, Jack and R{\'u}narsson, Kristj{\'a}n and Sadde, Shoval and Safari, Pegah and Sagot, Beno{\^{\i}}t and Sahala, Aleksi and Saleh, Shadi and Salomoni, Alessio and Samard{\v z}i{\'c}, Tanja and Samson, Stephanie and Sanguinetti, Manuela and San{\i}yar, Ezgi and S{\"a}rg, Dage and Saul{\={\i}}te, Baiba and Sawanakunanon, Yanin and Saxena, Shefali and Scannell, Kevin and Scarlata, Salvatore and Schneider, Nathan and Schuster, Sebastian and Schwartz, Lane and Seddah, Djam{\'e} and Seeker, Wolfgang and Seraji, Mojgan and Shahzadi, Syeda and Shen, Mo and Shimada, Atsuko and Shirasu, Hiroyuki and Shishkina, Yana and Shohibussirri, Muh and Sichinava, Dmitry and Siewert, Janine and Sigurðsson, Einar Freyr and Silveira, Aline and Silveira, Natalia and Simi, Maria and Simionescu, Radu and Simk{\'o}, Katalin and {\v S}imkov{\'a}, M{\'a}ria and Simov, Kiril and Skachedubova, Maria and Smith, Aaron and Soares-Bastos, Isabela and Sourov, Shafi and Spadine, Carolyn and Sprugnoli, Rachele and Steingr{\'{\i}}msson, Stein{\t h}{\'o}r and Stella, Antonio and Straka, Milan and Strickland, Emmett and Strnadov{\'a}, Jana and Suhr, Alane and Sulestio, Yogi Lesmana and Sulubacak, Umut and Suzuki, Shingo and Sz{\'a}nt{\'o}, Zsolt and Taguchi, Chihiro and Taji, Dima and Takahashi, Yuta and Tamburini, Fabio and Tan, Mary Ann C. and Tanaka, Takaaki and Tanaya, Dipta and Tella, Samson and Tellier, Isabelle and Testori, Marinella and Thomas, Guillaume and Torga, Liisi and Toska, Marsida and Trosterud, Trond and Trukhina, Anna and Tsarfaty, Reut and T{\"u}rk, Utku and Tyers, Francis and Uematsu, Sumire and Untilov, Roman and Ure{\v s}ov{\'a}, Zde{\v n}ka and Uria, Larraitz and Uszkoreit, Hans and Utka, Andrius and Vajjala, Sowmya and van der Goot, Rob and Vanhove, Martine and van Niekerk, Daniel and van Noord, Gertjan and Varga, Viktor and Villemonte de la Clergerie, Eric and Vincze, Veronika and Vlasova, Natalia and Wakasa, Aya and Wallenberg, Joel C. and Wallin, Lars and Walsh, Abigail and Wang, Jing Xian and Washington, Jonathan North and Wendt, Maximilan and Widmer, Paul and Wijono, Sri Hartati and Williams, Seyi and Wir{\'e}n, Mats and Wittern, Christian and Woldemariam, Tsegay and Wong, Tak-sum and Wr{\'o}blewska, Alina and Yako, Mary and Yamashita, Kayo and Yamazaki, Naoki and Yan, Chunxiao and Yasuoka, Koichi and Yavrumyan, Marat M. and Yenice, Arife Bet{\"u}l and Y{\i}ld{\i}z, Olcay Taner and Yu, Zhuoran and Yuliawati, Arlisa and {\v Z}abokrtsk{\'y}, Zden{\v e}k and Zahra, Shorouq and Zeldes, Amir and Zhou, He and Zhu, Hanzhi and Zhuravleva, Anna and Ziane, Rayan},
url = {http://hdl.handle.net/11234/1-4611},
note = {{LINDAT}/{CLARIAH}-{CZ} digital library at the Institute of Formal and Applied Linguistics ({{\'U}FAL})},
copyright = {Licence Universal Dependencies v2.9},
year = {2021}
}
| |||||
| Summarization | 10,126 articles | CC BY-NC-SA 4.0 | 2021 | ↗✓ 2026-06-28 | |
Professional article–summary pairs from BBC Bangla. Source: XL-Sum (Findings ACL 2021). @inproceedings{hasan-etal-2021-xl,
title = {{XL}-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages},
author = {Hasan, Tahmid and Bhattacharjee, Abhik and others},
booktitle = {Findings of ACL},
year = {2021}
} | |||||
| Text Classification | ~50k articles | CC BY-NC 4.0 | 2020 | ↗✓ 2026-06-11 | |
Annotated fake vs. authentic Bangla news articles. Source: BanFakeNews (LREC 2020). @inproceedings{hossain-etal-2020-banfakenews,
title = "{B}an{F}ake{N}ews: A Dataset for Detecting Fake News in {B}angla",
author = "Hossain, Md Zobaer and
Rahman, Md Ashraful and
Islam, Md Saiful and
Kar, Sudipta",
editor = "Calzolari, Nicoletta and
B{\'e}chet, Fr{\'e}d{\'e}ric and
Blache, Philippe and
Choukri, Khalid and
Cieri, Christopher and
Declerck, Thierry and
Goggi, Sara and
Isahara, Hitoshi and
Maegaard, Bente and
Mariani, Joseph and
Mazo, H{\'e}l{\`e}ne and
Moreno, Asuncion and
Odijk, Jan and
Piperidis, Stelios",
booktitle = "Proceedings of the Twelfth Language Resources and Evaluation Conference",
month = may,
year = "2020",
address = "Marseille, France",
publisher = "European Language Resources Association",
url = "https://aclanthology.org/2020.lrec-1.349/",
pages = "2862--2871",
language = "eng",
ISBN = "979-10-95546-34-4"
} | |||||
| Machine Translation | 2.75M pairs | Research only | 2020 | ↗✓ 2026-06-28 | |
High-quality Bangla–English parallel corpus built with aligner ensembling. Source: Not Low-Resource Anymore (EMNLP 2020). @inproceedings{hasan-etal-2020-low,
title = {Not Low-Resource Anymore: Aligner Ensembling, Batch Filtering, and New Datasets for {B}engali-{E}nglish Machine Translation},
author = {Hasan, Tahmid and Bhattacharjee, Abhik and others},
booktitle = {EMNLP},
year = {2020}
} | |||||
| Hate Speech Detection | 3,418 labelled statements | MIT | 2020 | ↗✓ 2026-09-08 | |
v1.0 release with one label per statement: geopolitical, religious, personal, gender abusive, or political hate. Repository license: MIT. Source: Bengali Hate Speech v1.0 (Karim et al., 2020). @misc{karim2020classificationbenchmarksunderresourcedbengali,
title={Classification Benchmarks for Under-resourced Bengali Language based on Multichannel Convolutional-LSTM Network},
author={Md. Rezaul Karim and Bharathi Raja Chakravarthi and John P. McCrae and Michael Cochez},
year={2020},
eprint={2004.07807},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2004.07807},
} | |||||
| Question Answering | 11,096 questions | Apache 2.0 | 2020 | ↗✓ 2026-09-08 | |
Bengali primary-task data with 10,768 train and 328 dev questions for passage selection and minimal answer extraction; not the GoldP subset. Source: TyDi QA primary tasks v1.0 (TACL 2020). @article{clark-etal-2020-tydi,
title = "{T}y{D}i {QA}: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages",
author = "Clark, Jonathan H. and
Choi, Eunsol and
Collins, Michael and
Garrette, Dan and
Kwiatkowski, Tom and
Nikolaev, Vitaly and
Palomaki, Jennimaria",
editor = "Johnson, Mark and
Roark, Brian and
Nenkova, Ani",
journal = "Transactions of the Association for Computational Linguistics",
volume = "8",
year = "2020",
address = "Cambridge, MA",
publisher = "MIT Press",
url = "https://aclanthology.org/2020.tacl-1.30/",
doi = "10.1162/tacl_a_00317",
pages = "454--470"
} | |||||
| Named Entity Recognition | 12,000 sentences | ODC-By; research use only | 2019 | ↗✓ 2026-09-08 | |
Silver standard Wikipedia PER/LOC/ORG annotations in the Rahimi et al. balanced splits; Bengali train/validation/test contain 10,000/1,000/1,000 sentences. The license field records the original release's ODC-By and 'For research use only' statements. Source: Massively Multilingual Transfer for NER (ACL 2019). @inproceedings{rahimi-etal-2019-massively,
title = "Massively Multilingual Transfer for {NER}",
author = "Rahimi, Afshin and
Li, Yuan and
Cohn, Trevor",
editor = "Korhonen, Anna and
Traum, David and
M{\`a}rquez, Llu{\'i}s",
booktitle = "Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics",
month = jul,
year = "2019",
address = "Florence, Italy",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/P19-1015/",
doi = "10.18653/v1/P19-1015",
pages = "151--164"
} | |||||
| Sentiment Analysis | 5,038 comments | Research only | 2018 | ↗✓ 2026-04-02 | |
Aspect-based sentiment in two domains. Source: Aspect-based sentiment (cricket, restaurant). @article{Rahman_2018, title={Datasets for Aspect-Based Sentiment Analysis in Bangla and Its Baseline Evaluation}, volume={3}, ISSN={2306-5729}, url={http://dx.doi.org/10.3390/data3020015}, DOI={10.3390/data3020015}, number={2}, journal={Data}, publisher={MDPI AG}, author={Rahman, Md. Atikur and Kumar Dey, Emon}, year={2018}, month=May, pages={15} } | |||||
| Text Classification | 376,226 articles | Research only | 2018 | ↗✓ 2026-02-14 | |
News articles across 5 categories. Source: Bangla Article Dataset. @inproceedings{Tanvir_Alam_2018, title={BARD: Bangla Article Classification Using a New Comprehensive Dataset}, url={http://dx.doi.org/10.1109/ICBSLP.2018.8554382}, DOI={10.1109/icbslp.2018.8554382}, booktitle={2018 International Conference on Bangla Speech and Language Processing (ICBSLP)}, publisher={IEEE}, author={Tanvir Alam, Md and Mofijul Islam, Md}, year={2018}, month=Sept, pages={1–5} } | |||||
| ASR / TTS | 196k utterances / 229 hrs | CC BY-SA 4.0 | 2018 | ↗✓ 2026-06-28 | |
Crowdsourced Bangladeshi Bangla read speech. Source: Google crowdsourced speech. @inproceedings{Kjartansson_2018, series={sltu_2018}, title={Crowd-Sourced Speech Corpora for Javanese, Sundanese, Sinhala, Nepali, and Bangladeshi Bengali}, url={http://dx.doi.org/10.21437/SLTU.2018-11}, DOI={10.21437/sltu.2018-11}, booktitle={6th Workshop on Spoken Language Technologies for Under-Resourced Languages (SLTU 2018)}, publisher={ISCA}, author={Kjartansson, Oddur and Sarin, Supheakmungkol and Pipatsrisawat, Knot and Jansche, Martin and Ha, Linne}, year={2018}, month=Aug, pages={52–55}, collection={sltu_2018} } | |||||
| POS Tagging | ~7,390 sentences | Research only | 2010 | ↗✓ 2026-07-20 | |
Classic manually tagged corpus, 40-tag tagset. Source: Society for Natural Language Technology Research. No BibTeX on file for this dataset. Rather than generate one, we leave it out — add the published entry via PR. | |||||
Nothing matches those filters. Add a missing dataset via PR.