Researcher · Multimodal Learning & Vision-Language

Rabeya Akter

I am a Research Assistant in the Department of Electrical and Computer Engineering at North South University, where I work with Md Salman Shamil and Dr. Md Adnan Arefeen on egocentric video understanding, modeling hand-object interactions and everyday task actions with LLM-guided reasoning.

My latest work, OPASD, teaches reasoning language models where to look. This work is supervised by Dr. Md Mofijul Islam (Amazon GenAI, USA) and Dr. Tasnim Mohiuddin (Qatar Computing Research Institute, Qatar). I am also working with Dr. Ser-Nam Lim (University of Central Florida, USA) on world models. My first-author work, EK-Bangla, asks whether multimodal LLMs can describe fine-grained hand-object interactions in Bangla as well as they do in English.

I earned my B.Sc. in Robotics and Mechatronics Engineering from the University of Dhaka. After graduating, I worked on amodal counting with Dr. Md Mehedi Hasan (University of Dhaka). As an undergraduate, I worked on continuous Bangla sign language translation with Dr. Shafin Rahman (North South University) and Dr. Sejuti Rahman. In industry, I was an Associate Software Engineer on the machine learning team at Therap (BD) Ltd. and a data science intern at Pathao.

Beyond research, I enjoy mentoring in olympiads, taking part in hackathons, and exploring cultural activities such as writing and performing.

I am actively seeking PhD opportunities for Fall 2027.

Research interests

  • Language Model Post-Training
  • Reinforcement Learning
  • Multimodal Learning
  • Vision-Language Models
  • Embodied AI
  • Computer Vision

Research

Most agents learn from outcomes, which tell them whether they succeeded but not why. My work gives an agent richer signals across the three abilities it needs. To reason, it learns where to look, not just which answer is right. To perceive, it fuses modalities so that one view fills in what another misses. To act, it learns why it failed, as a vision–language model turns its mistakes into rewards.

01 · Reason

Reinforcement Learning & Post-Training

I teach language models which earlier steps to focus on, so they reason more accurately with shorter answers.

02 · Perceive

Multimodal Learning

I fuse complementary modalities, like video with body pose or time series with text, so models capture what any single view would miss.

03 · Act

Embodied AI

I use vision–language models to explain why a robot failed and turn that feedback into rewards.

News

Publications & Preprints

Google Scholar →
Preprint 2026 Under review

EK-Bangla: When Hand-Object Interaction Meets Low-Resource Generation

Rabeya Akter, Md Salman Shamil, A. K. M. Fazley Rabbi, Arpa Deb, Ridwain Islam, Md Adnan Arefeen

EK-Bangla is the first benchmark for fine-grained egocentric hand-object interaction in Bangla, built on EPIC-KITCHENS-100: 9,600 native-speaker-verified captions and 2,000 four-way MCQs. Across nine MLLMs, rankings flip between captioning and MCQ, and every model scores lower in Bangla than in English on identical clips.

Project page Paper & code coming soon
ICML 2026

LAGEA: Language Guided Embodied Agents for Robotic Manipulation

Abdul Monaf Chowdhury, Akm Moshiur Rahman Mazumder, Safaeid Hossain Arib, Rabeya Akter

LAGEA turns structured vision-language-model reflections on a robot's own failures into temporally grounded, progress-based shaping rewards. It localizes the decisive moments of an episode, aligns feedback with visual states, and fades the guidance as the policy improves, outperforming FuRL on Meta-World MT10 and Fetch.

AAAI 2026

T3Time: Tri-Modal Time Series Forecasting via Adaptive Multi-Head Alignment and Residual Fusion

Abdul Monaf Chowdhury, Rabeya Akter, Safaeid Hossain Arib

T3Time reads a time series three ways (temporal, spectral, and through an LLM prompt), gates temporal versus spectral features by forecast horizon, and fuses modalities with adaptively weighted cross-modal attention heads. It sets the best MSE on 7 of 8 long-term benchmarks and stays strong with only 5–10% of the training data.

arXiv 2025 Under review

Counting Through Occlusion: Framework for Open World Amodal Counting

Safaeid Hossain Arib, Rabeya Akter, Abdul Monaf Chowdhury, Md Jubair Ahmed Sourov, Md Mehedi Hasan

State-of-the-art counters only count what they can see. CountOCC reconstructs the features of hidden objects from visible context and text/exemplar priors, and enforces attention consistency between occluded and unoccluded views. It sets a new state of the art on our occlusion benchmarks FSC-147-OCC and CARPK-OCC, and on CAPTURe-Real.

PLOS ONE 2025

SignFormer-GCN: Continuous Sign Language Translation Using Spatio-Temporal Graph Convolutional Networks

Safaeid Hossain Arib, Rabeya Akter, Sejuti Rahman, Shafin Rahman

SignFormer-GCN jointly encodes RGB video with a transformer and skeletal keypoints with an STGCN-LSTM, capturing both broad context and fine-grained body motion. It delivers competitive gloss-free translation on German, American, and Bangla sign language benchmarks with only 9.43M parameters.

arXiv 2023

Bornil: An Open-Source Sign Language Data Crowdsourcing Platform for AI Enabled Dialect-Agnostic Communication

Shahriar Elahi Dhruvo, Mohammad Akhlaqur Rahman, Manash Kumar Mandal, Md. Istiak Hossain Shihab, A. A. Noman Ansary, Kaneez Fatema Shithi, Sanjida Khanom, Rabeya Akter, Safaeid Hossain Arib, M.N. Ansary, Sazia Mehnaz, Rezwana Sultana, Sejuti Rahman, Sayma Sultana Chowdhury, Sabbir Ahmed Chowdhury, Farig Sadeque, Asif Sushmit

Bornil is an open-source platform for recording, annotating, and validating multilingual sign language data. It was used to build BornilDB v1.0, the largest Bangladeshi Sign Language dataset (73 hours, 21,000+ samples), with recognition benchmarks for this low-resource setting.

Experience

Research

  1. Mar 2026 – Present

    Research Assistant

    Dept. of Electrical and Computer Engineering, North South University · with Md Salman Shamil and Dr. Md Adnan Arefeen

    • Working on egocentric video understanding, modeling hand-object interactions and daily task actions with LLM-guided reasoning.
    • Led EK-Bangla, the first Bangla benchmark for fine-grained egocentric hand-object interaction, with 9,600 verified captions and 2,000 MCQs. Submitted to WACV 2027.
  2. Jun 2025 – Feb 2026

    Research Assistant

    Dept. of Robotics and Mechatronics Engineering, University of Dhaka · with Dr. Md Mehedi Hasan and Md. Jubair Ahmed Sourov

    • Designed and implemented CountOCC, the first open-world amodal counting framework, and assembled occlusion-augmented datasets to evaluate it.
    • Co-authored the resulting manuscript, submitted to WACV 2027.
  3. Jan 2023 – Dec 2023

    Undergraduate Research Assistant

    Dept. of Robotics and Mechatronics Engineering, University of Dhaka · with Dr. Shafin Rahman and Dr. Sejuti Rahman

    • Designed and implemented a novel architecture for continuous sign language translation that achieved competitive performance across Bangla, American, and German benchmarks.
    • Co-authored the resulting work, SignFormer-GCN, published in PLOS ONE.
    • Secured a research grant from the Bangladesh Ministry of Science and Technology's Special Innovation Fund (Project ID: SRG-232431, FY 2023–2024).

Industry

  1. Apr 2024 – May 2025

    Associate Software Engineer, QA

    Therap (BD) Ltd. · Machine Learning team · Dhaka, Bangladesh

    • Researched Amazon Alexa devices and sound detection systems, reviewed action recognition datasets, and contributed to the design of data collection for action recognition.
    • Led quality assurance for the video redaction system, building test plans and cases to validate the machine learning algorithms used in video processing.
  2. Jan 2024 – Mar 2024

    AIM Intern

    Pathao Limited · Data Science, Analytics & Insights · Dhaka, Bangladesh

    • Developed five KPIs and metrics for measuring the brand health of Pathao Courier, aligned with company-wide strategic goals.
    • Designed data-driven questionnaires to assess brand health, improving response quality and depth of insights.

Selected Projects