I'm an AI researcher working at the intersection of speech and large language models. For the past three years I built production speech systems at Decodis while publishing from the Apurba–NSU R&D Lab — deliberately, because I wanted the research to survive contact with real audio.
The next billion people will talk to machines, not type — and today's models barely hear them.
The problem I'm becoming an expert in
Speech LLMs are converging on a fluent, low-latency, full-duplex interface — but almost entirely in English, under clean conditions. The failure is not a data shortage that more hours will fix. It is structural. Tokenisation schemes discard the morphology that agglutinative languages carry meaning in. Aggregate error rates hide exactly the code-switching and suffix-level failures that matter. End-to-end systems align adapters and embeddings without the text-level supervision a jointly speaking model needs. I want to build a full-duplex speech-to-speech model that learns from real, noisy, multilingual conversation — retrieval-grounded so it can be held to what is true, trained through shared latent representations rather than bolted-together modules, and evaluated on metrics that expose where it breaks instead of averaging over it.
Evidence I've already made progress on it
- I did not trust the metric, so I replaced it. Aggregate WER hides code-switching failure, so I built CS-WER to measure suffix-, switch-, and root-level performance separately — then a dynamic block-online model that cut errors 17.5% relative to conventional streaming. Interspeech 2026, oral.
- The systems work on real audio, not benchmarks. Production ASR across four languages at ~100k-audio scale; Swahili word error rate 60% → 28% on noisy overlapped field recordings. Systems in Production.
- Deployment is a research constraint, not an afterthought. A distillation framework removing expensive hyper-parameter search, generalising across classification, data-free distillation, and detection. PLOS ONE.
- The work reaches people. A voice-first Bangla health assistant for low-literacy users, built with public-health experts at icddr,b — $30K funded, and I served as Co-PI. MenoChat.
Highlights
-
Apr 2026Selected for the Adaption Labs inaugural grant cohort (AdaptBench, LLM evaluation) — the adaptive-AI lab founded by Sara Hooker. -
2026First-author streaming ASR paper accepted at Interspeech 2026 — invited for oral presentation. -
2026Point Cloud Transformer for physical rehabilitation at IEEE JBHI. -
USPTO
Aug 2026U.S. patent application filed: Theme-Based Voice Emotion Analytics Method (#19/763,175).
-
decodis
Jan 2026Concluded two and a half years at Decodis as a Machine Learning Engineer.
-
PRL
2024DeepMarkerNet published in Pattern Recognition Letters (Elsevier).
-
2024Low-resource Bangla TTS presented at PML4LRS @ ICLR 2024. -
2023Knowledge-distillation paper published in PLOS ONE; later featured at Cohere Labs. -
decodis
Jul 2023Joined Decodis (Boston, remote) as a Machine Learning Engineer.
Systems in Production
At Decodis I shipped ASR for Swahili, Yoruba, Azerbaijani, and Hausa into workflows running at roughly 100k audios. On real noisy field audio with heavy speaker overlap, Swahili word error rate went from 60% to 28% and Azerbaijani from 36% to 18% — achieved by attacking the high-error tail rather than the average, with SNR-based filtering, controlled-ratio overlap augmentation, and pitch-aligned speaker concatenation feeding a recursive training loop.
▶ Hear it — Swahili field audio
Publications
▶ See it — Bangla–English code-switching
Work in Progress
Patents
Selected Projects
MenoChat — Voice-First Women's Health AssistantExperience
- Built production ASR for Swahili, Yoruba, Azerbaijani, and Hausa, integrated into end-to-end workflows operating at roughly 100k audios at scale.
- Reduced Swahili WER from 60% to 28% on a highly noisy, overlapped test set via a voting-based SNR filter, controlled-ratio overlap augmentation, pitch-aligned speaker concatenation, and a recursive training loop.
- Cut Azerbaijani WER on noisy field audio from 36% to 18% through percentile-based error analysis targeting the high-error tail rather than the average.
- Designed the cross-language timestamping approach that became the basis of a pending patent; added BERTScore evaluation that surfaced mislabeled training data.
- Designed and led the company's data-annotation process end to end; delivered $25,000+ in cost savings per project across three projects.
- Lead researcher on speech LLMs, ASR, TTS, knowledge distillation, and agentic AI; published at Interspeech, ICLR, PLOS ONE, PRL, and IEEE JBHI.
- Appointed Co-Principal Investigator on the funded MenoChat project.
- Mentored junior researchers, carrying projects from raw data to publication; that mentorship contributed to work at Pattern Recognition Letters and the ICLR PML4LRS workshop.
Collaborators
- Dr. Nabeel Mohammed — supervisor, Apurba–NSU. Co-author, Interspeech & JBHI. Scholar
- Dr. Shafin Rahman — co-supervisor. Co-author, JBHI & distillation. Scholar
- Dr. Fuad Rahman — Apurba Technologies. Co-author across speech & distillation. Scholar
- Pravarakya Reddy Battula — CPTO, Decodis. Supervisor for multilingual ASR & the pending patent.
Teaching
- Edge Bangladesh — ML & Deep Learning for Industry Engineers — designed and taught a three-month course for ~20 practising engineers at MIR Group; authored the full curriculum, from ML foundations to Transformer attention.
- Hands-on PyTorch and Research — workshop session on signal processing and vision methods, North South University, Fall 2024.
- Mentorship — junior lab members at Apurba–NSU, from raw data to publications at PRL and the ICLR PML4LRS workshop.
Education
Honors & Grants
- $30,000 Research Grant + Best Innovative Idea + Best Poster — NCSRHR 2023 (MenoChat)
- Adaption Labs Inaugural Grant Cohort — 2026 (AdaptBench)
- Research featured at Cohere Labs
(Cohere For AI) — knowledge-distillation work from PLOS ONE
- CTRG 2023 Research Grant — 5 Lac BDT (~$4,500), Bangladesh national research grant
- NAACL Mentorship Program — selected mentee, low-resource virtual assistant project
- Summa Cum Laude — North South University
- Merit-Based Scholarship — North South University, nationwide competitive admission (2019)
- The Daily Star Award — outstanding O-Levels (2016) and A-Levels (2018)
Contact
I like meeting people who care about speech, multilingual AI, or health. Reach me at kingrafat82@gmail.com.
Download CV
. Integrating Bangla ASR and TTS had stalled others; the end-to-end voice pipeline was made to work here. Prototype scored 75 on an accessibility-adapted SUS.