01 · Perceive
Multimodal Learning
I combine vision, language, pose and time series so models see what a single view misses, and test whether that understanding holds in low-resource languages like Bangla.
I am a Research Assistant at North South University, working with Md Salman Shamil and Dr. Md Adnan Arefeen on egocentric video understanding (EK-Bangla).
My latest work, OPASD, teaches reasoning language models where to look. This work is supervised by Dr. Md Mofijul Islam (Amazon GenAI, USA) and Dr. Tasnim Mohiuddin (Qatar Computing Research Institute, Qatar). I am also working with Dr. Ser-Nam Lim (University of Central Florida, USA) on world models.
I earned my B.Sc. in Robotics and Mechatronics Engineering from the University of Dhaka, where I worked on amodal counting with Dr. Md Mehedi Hasan and on sign language translation with Dr. Shafin Rahman and Dr. Sejuti Rahman.
Outside of research, I love good books and well-written blogs, and I am always happy to chat about either, so please feel free to reach out.
I am actively seeking PhD opportunities for Fall 2027.
Research interests
Most agents learn from outcomes, which tell them whether they succeeded but not why. My work gives an agent richer signals across the three abilities it needs. To perceive, it fuses modalities so that one view fills in what another misses. To reason, it learns where to look, not just which answer is right. To act, it learns why it failed, as a vision–language model turns its mistakes into rewards.
01 · Perceive
I combine vision, language, pose and time series so models see what a single view misses, and test whether that understanding holds in low-resource languages like Bangla.
02 · Reason
I teach language models which earlier steps to focus on, so they reason more accurately with shorter answers.
03 · Act
I use vision–language models to explain why a robot failed and turn that feedback into rewards.
New Our new preprint, Teach Yourself Where to Look, is now on arXiv.
New Introducing EK-Bangla, my first-author benchmark for fine-grained egocentric hand-object interaction in Bangla. The project page is live.
Our paper LAGEA: Language Guided Embodied Agents for Robotic Manipulation has been accepted at ICML 2026.
Joined the Dept. of Electrical and Computer Engineering, North South University, as a Research Assistant, working on egocentric video understanding.
Our work T3Time has been accepted at the AAAI 2026 Main Technical Track.
SignFormer-GCN has been accepted at the WiML Workshop at NeurIPS 2025.
Started as a Research Assistant at the Dept. of Robotics and Mechatronics Engineering, University of Dhaka, working on amodal counting.
SignFormer-GCN is officially published in PLOS ONE.
Joined Therap (BD) Ltd. as an Associate Software Engineer on the Machine Learning team.
Graduated from the Department of Robotics and Mechatronics Engineering, University of Dhaka.
Successfully defended my undergraduate thesis — a huge milestone in my academic journey.
Presented our poster Classical machine learning approach for human activity recognition using location data at UbiComp/ISWC 2021.
On-policy self-distillation tells a student what to predict, but not where to look. OPASD adds solution-conditioned attention distillation, projecting a privileged teacher's attention onto student-visible context. It beats token-only OPSD by 4.98–8.40 points across three Qwen3 sizes while cutting rollout tokens by 73.9%.
EK-Bangla is the first benchmark for fine-grained egocentric hand-object interaction in Bangla, built on EPIC-KITCHENS-100: 9,600 native-speaker-verified captions and 2,000 four-way MCQs. Across nine MLLMs, rankings flip between captioning and MCQ, and every model scores lower in Bangla than in English on identical clips.
LAGEA turns structured vision-language-model reflections on a robot's own failures into temporally grounded, progress-based shaping rewards. It localizes the decisive moments of an episode, aligns feedback with visual states, and fades the guidance as the policy improves, outperforming FuRL on Meta-World MT10 and Fetch.
T3Time reads a time series three ways (temporal, spectral, and through an LLM prompt), gates temporal versus spectral features by forecast horizon, and fuses modalities with adaptively weighted cross-modal attention heads. It sets the best MSE on 7 of 8 long-term benchmarks and stays strong with only 5–10% of the training data.
State-of-the-art counters only count what they can see. CountOCC reconstructs the features of hidden objects from visible context and text/exemplar priors, and enforces attention consistency between occluded and unoccluded views. It sets a new state of the art on our occlusion benchmarks FSC-147-OCC and CARPK-OCC, and on CAPTURe-Real.
SignFormer-GCN jointly encodes RGB video with a transformer and skeletal keypoints with an STGCN-LSTM, capturing both broad context and fine-grained body motion. It delivers competitive gloss-free translation on German, American, and Bangla sign language benchmarks with only 9.43M parameters.
Bornil is an open-source platform for recording, annotating, and validating multilingual sign language data. It was used to build BornilDB v1.0, the largest Bangladeshi Sign Language dataset (73 hours, 21,000+ samples), with recognition benchmarks for this low-resource setting.
Our entry to the Sussex-Huawei Locomotion Challenge 2021 recognizes eight locomotion and transportation activities from smartphone location data using statistical features and a random forest, reaching 78.14% validation accuracy and a 78.28% weighted F1 score.
Dept. of Electrical and Computer Engineering, North South University · with Md Salman Shamil and Dr. Md Adnan Arefeen
Dept. of Robotics and Mechatronics Engineering, University of Dhaka · with Dr. Md Mehedi Hasan and Md. Jubair Ahmed Sourov
Dept. of Robotics and Mechatronics Engineering, University of Dhaka · with Dr. Shafin Rahman and Dr. Sejuti Rahman
Therap (BD) Ltd. · Machine Learning team · Dhaka, Bangladesh
Pathao Limited · Data Science, Analytics & Insights · Dhaka, Bangladesh
Course project · Aug – Nov 2022
A reinforcement learning agent that learns when to buy, sell, or hold three Bangladeshi stocks, using Q-learning with linear function approximation, tested with and without the COVID-19 period.
View project →
Course project · Aug – Nov 2023
Can a robot comfort people as well as a human? We programmed NAO to respond to happiness, sadness, and fear, then compared it with human support on emotional impact, comfort, clarity, and satisfaction.
View project →