Kazi Rafat Haa Meem
Kazi Rafat Haa Meem
Speech & Language AI Researcher
Apurba–NSU R&D Lab · Dhaka, Bangladesh
Publishes as Kazi Rafat · also cited as K. Rafat
Summa Cum Laude · North South University Merit-Based Scholarship · nationwide competitive admission $34.5K+ in research grants awarded Adaption Labs inaugural cohort

I'm an AI researcher working at the intersection of speech and large language models. For the past three years I built production speech systems at Decodis while publishing from the Apurba–NSU R&D Lab — deliberately, because I wanted the research to survive contact with real audio.

The next billion people will talk to machines, not type — and today's models barely hear them.

The problem I'm becoming an expert in

Speech LLMs are converging on a fluent, low-latency, full-duplex interface — but almost entirely in English, under clean conditions. The failure is not a data shortage that more hours will fix. It is structural. Tokenisation schemes discard the morphology that agglutinative languages carry meaning in. Aggregate error rates hide exactly the code-switching and suffix-level failures that matter. End-to-end systems align adapters and embeddings without the text-level supervision a jointly speaking model needs. I want to build a full-duplex speech-to-speech model that learns from real, noisy, multilingual conversation — retrieval-grounded so it can be held to what is true, trained through shared latent representations rather than bolted-together modules, and evaluated on metrics that expose where it breaks instead of averaging over it.

Evidence I've already made progress on it

Speech LLMsSpeech RecognitionSpeech-to-SpeechText-to-SpeechAudio Signal ProcessingAgentic AIKnowledge DistillationModel CompressionHealth AILow-Resource & Code-Switching
In progress Building end-to-end audio language models and full-duplex speech-to-speech systems. Applying to PhD programs.

Highlights

scroll for more ↓

Systems in Production

At Decodis I shipped ASR for Swahili, Yoruba, Azerbaijani, and Hausa into workflows running at roughly 100k audios. On real noisy field audio with heavy speaker overlap, Swahili word error rate went from 60% to 28% and Azerbaijani from 36% to 18% — achieved by attacking the high-error tail rather than the average, with SNR-based filtering, controlled-ratio overlap augmentation, and pitch-aligned speaker concatenation feeding a recursive training loop.

▶ Hear it — Swahili field audio
ReferenceWhatsApp yangu mimi hutumia kuongea na watu walio nje ya nchi, na kazi zangu pia nyingi ni ya WhatsApp na marafiki pia huwa tunaongea kupitia WhatsApp
Model outputwhatsapp yangu mimi hutumia kuongea na watu walio nje yaaa nchi, na kazi zangu pia nyingi ni ya whatsapp na marafika pia huwa tunaongea kupitia whatsapp
Conversational Swahili from the deployed pipeline. Remaining errors are underlined.

Publications

Dynamic Block-Online Streaming ASR for Low-Resource Agglutinative Code-Switching Speech with Morphology-Aware Evaluation
Interspeech 2026OralFirst author
K. Rafat, A. Imran, M. I. Hossain, M. R. Ali, F. Rahman, S. Rahman, N. Mohammed
A dynamic block-online model trained with global attention, achieving a 17.5% relative error reduction over conventional streaming approaches on code-switched speech. Introduces CS-WER, an evaluation framework measuring suffix-, switch-, and root-level performance separately.
▶ See it — Bangla–English code-switching
Referencecollegeটি ঢাকা শিক্ষা boardের আওতাধীন এবং জাতীয় বিশ্ববিদ্যালয়ের অধিভক্ত
Ours — dynamic block-onlinecollegeটি ঢাকা শিক্ষা bordের আওতাধীন এবং জাতীয় বিশ্ববিদ্যালয়ের অধিভক্ত
Conventional streamingcollegeটি ঢা শিক্ষা bordর আওতাধীন এবং জাতীয় বিশ্ববিদ্যালয়ের অধিভক্ত
Conventional streaming truncates at block boundaries (marked): ঢাকা clipped to ঢা, and the Bangla suffix ের on the English loanword board reduced to র. Exactly the failure aggregate WER hides — and what CS-WER was built to measure.
A Point Cloud Transformer for Remote Monitoring and Automated Assessment of Physical Rehabilitation Exercises
IEEE JBHI 2026IF 7.7 · Q1First author
K. Rafat, Md. I. Hossain, M. M. L. Elahi, S. Momen, F. Rahman, N. Mohammed, S. Rahman
Geometric transformer with low-cost attention that classifies and scores rehabilitation exercise quality from 3D pose and motion data. Supported by the CTRG 2023 national grant.
Mitigating Carbon Footprint of Hyper-parameter Selection During Knowledge Distillation
PLOS ONE 2023Q1First author
K. Rafat, A. A. Mahfug, M. I. Hossain, S. Momen, F. Rahman, S. Rahman, N. Mohammed
A distillation framework that removes computationally expensive hyper-parameter search, generalising across image classification, data-free distillation, and object detection. Began as an undergraduate course project.
DeepMarkerNet: Supervised Duchenne Marker Modeling for Spontaneous Smile Recognition
Pattern Recognition Letters 2024IF 3.9Co-author
M. J. Hasan, K. Rafat, F. Rahman, N. Mohammed, S. Rahman · Elsevier
Precision-Driven Low-Resource Speech Synthesis for Bangla Text-to-Speech System
PML4LRS @ ICLR 2024Co-author
T. S. Shahjahan, M. I. Hossain, K. Rafat, M. R. Amin, F. Rahman, N. Mohammed

Work in Progress

Latent-RAG: End-to-End Speech-to-Speech Modelling with Retrieval Grounding and Controllable GenerationIn preparation
A low-latency, barge-in-capable full-duplex speech-to-speech system for low-resource languages. Built around shared latent representations so components train genuinely end to end rather than as independently optimised modules, with training objectives chosen for text-level supervision in a jointly speaking system.

Patents

Theme-Based Voice Emotion Analytics MethodPatent Pending
U.S. Patent Application #19/763,175 · Decodis · 2026
An analytics method augmenting speech-based prosodic information, including a timestamping approach for accurate grapheme and prosodic alignment in real-time multilingual transcription.

Selected Projects

MenoChatMenoChat — Voice-First Women's Health Assistant
$30K GrantBest Idea & Best PosterCo-Principal Investigator
An agentic health assistant for women with limited literacy in Bangladesh, combining Bangla ASR and TTS with a multi-stage, comorbidity-aware LLM pipeline that grounds answers in verified medical information and escalates high-risk cases. Refined with public-health experts at icddr,b icddr,b. Integrating Bangla ASR and TTS had stalled others; the end-to-end voice pipeline was made to work here. Prototype scored 75 on an accessibility-adapted SUS.
AdaptBench — Benchmark for LLM Evaluation
Adaption Labs 2026
Selected for the inaugural Adaption Labs grant cohort to build a benchmark for evaluating large language models.
Intelligent Virtual Assistant for Low-Resource Devices
NAACL Mentorship Program
Multilingual, multimodal voice assistant on constrained hardware using distilled language models, with mood recognition and parental controls — reasoning about model compression, multilingual coverage, and on-device deployment simultaneously.

Experience

DecodisJul 2023 – Jan 2026
Machine Learning Engineer · Boston, MA (Remote)
  • Built production ASR for Swahili, Yoruba, Azerbaijani, and Hausa, integrated into end-to-end workflows operating at roughly 100k audios at scale.
  • Reduced Swahili WER from 60% to 28% on a highly noisy, overlapped test set via a voting-based SNR filter, controlled-ratio overlap augmentation, pitch-aligned speaker concatenation, and a recursive training loop.
  • Cut Azerbaijani WER on noisy field audio from 36% to 18% through percentile-based error analysis targeting the high-error tail rather than the average.
  • Designed the cross-language timestamping approach that became the basis of a pending patent; added BERTScore evaluation that surfaced mislabeled training data.
  • Designed and led the company's data-annotation process end to end; delivered $25,000+ in cost savings per project across three projects.
Apurba Technologies · NSU R&D Lab2022 – Present
Research Assistant · Dhaka, Bangladesh · concurrent with Decodis, 2023–2026
  • Lead researcher on speech LLMs, ASR, TTS, knowledge distillation, and agentic AI; published at Interspeech, ICLR, PLOS ONE, PRL, and IEEE JBHI.
  • Appointed Co-Principal Investigator on the funded MenoChat project.
  • Mentored junior researchers, carrying projects from raw data to publication; that mentorship contributed to work at Pattern Recognition Letters and the ICLR PML4LRS workshop.

Collaborators

Teaching

Education

North South UniversityJan 2019 – May 2022
B.Sc. Computer Science & Engineering · Summa Cum Laude
Merit-Based Scholarship throughout (nationwide competitive admission). Core coursework in Pattern Recognition & Neural Networks, Machine Learning, and Probability & Statistics.

Honors & Grants

Contact

I like meeting people who care about speech, multilingual AI, or health. Reach me at kingrafat82@gmail.com.

Download CV