EK-Bangla: When Hand-Object Interaction Meets Low-Resource Generation
Under review at WACV 2027 (Round 2)
TL;DR
Same video, same model, different language: multimodal LLMs that describe hand-object interactions well in English drop hand, object, and verb details in Bangla. EK-Bangla measures that gap with a caption track and a multiple-choice track built from the same egocentric clips.
- 9,600human-verified Bangla captions in EK-BanglaCAP
- 2,000four-way MCQs in EK-BanglaMCQ, from 1,346 verified captions
- 54.5%of MLLM draft captions needed native-speaker revision
- up to −23.9pt accuracy drop for Bangla vs. English MCQs

Abstract
Describing hand-object interactions in egocentric video requires resolving the acting hand, manipulated object, contact, and temporal progression details important for user-facing visual systems. Although multimodal large language models (MLLMs) are increasingly evaluated on fine-grained egocentric understanding, their ability to preserve such grounding in low-resource languages remains unclear. We introduce EK-Bangla, a two-track Bangla benchmark built on the EPIC-KITCHENS-100 validation split. EK-BanglaCAP contains 9,600 MLLM-generated Bangla captions verified by native speakers under a nine-category error taxonomy, with 54.5% requiring revision. EK-BanglaMCQ contains 2,000 four-way questions constructed from verified captions through a proposer, a distractor, and a verifier pipeline that rejects 76.07% of candidates, followed by human validation for quality assurance. Across nine open- and closed-source MLLMs, performance varies substantially across captioning and multiple-choice questions: Qwen3.5-9B, for example, trails GPT-5 mini by 17.44% in chrF++ on captioning yet exceeds it by 7.20% accuracy on MCQ. Under matched visual evidence, all nine models perform worse on Bangla than English MCQs, with drops up to 23.90%. These results show that fine-grained multimodal capability does not transfer uniformly across language and evaluation format, motivating diagnostic evaluation of low-resource egocentric understanding.
Fine-grained, but only in English
Motivation
A user-facing visual assistant has to say which hand did what to which object, and in what order, and it has to say it in the user's own language. Egocentric video makes this hard: the evidence often occupies a small part of the frame, hands and objects occlude each other, the camera moves, and the meaning of an action depends on how it unfolds over time.
Fine-grained hand-object benchmarks such as EgoHOIBench and HanDyVQA are English-only, while multilingual and Bangla benchmarks focus on images or coarse scene-level video. Nothing tests whether an MLLM can convey subtle interaction details precisely in a low-resource language. EK-Bangla fills that gap for Bangla.
Two tracks, one set of clips
Benchmark construction

EK-BanglaCAP: free-form captions
- DraftEight frames sampled at full 1920×1080 resolution, together with the EPIC narration, verb and noun as weak context, go to Gemini 3.1 Pro. A first pass records hand roles, contact, motion and action phases. A second pass rewrites this into one fluent Bangla sentence.
- VerifyNative Bangla speakers watch every source clip and revise the caption where needed. 5,234 of 9,600 drafts (54.5%) needed correction.
- Label errorsEach corrected draft is tagged with a nine-category taxonomy: six visual errors (wrong verb, object, hand or tool, hallucination, omission) and three linguistic ones (unnatural wording, untranslated text, grammar or spelling).
Verified captions average 24.7 words, against 2.9 words for the original EPIC narrations.
EK-BanglaMCQ: multiple choice
- ProposeWorking from caption text alone, so every answer traces back to human-verified content, a proposer picks the applicable categories and writes one question per category.
- DistractA distractor generator writes three plausible but wrong options that match the correct answer in form, detail and length.
- VerifyA verifier sees only the question and shuffled options, with no video. Any item it answers correctly is rejected, so questions solvable from priors or lexical cues are filtered out. 6,358 of 8,358 proposals (76.07%) were rejected.
A 300-item human audit (15% of the benchmark), with each item judged by two of three annotators, approved 100% of questions as valid, 100% of answers as correct, and 98% of distractors as plausible.
| Category | What it tests | Questions |
|---|---|---|
| Action | Manipulation performed by the hands | 332 |
| Object | Object or tool being acted on | 780 |
| Hand-Object Relation | Which hand contacts which object | 206 |
| Temporal | Order in which events occur | 204 |
| Location | Source or destination of the action | 454 |
| State | Observable change in an object | 24 |

Results
Nine MLLMs · 8 frames at 456×256
We evaluate three closed-source models (GPT-5.1, GPT-5 mini, Gemini-2.5-Flash) and six open-source models from 4B to 30B parameters.
Captioning: no single leader
| Model | BLEU-4 | ROUGE-L | METEOR | CIDEr | chrF++ |
|---|---|---|---|---|---|
| GPT-5.1 | 5.77 | 28.02 | 30.90 | 12.75 | 35.58 |
| GPT-5 mini | 3.34 | 23.25 | 23.64 | 9.46 | 30.85 |
| Gemini-2.5-Flash | 9.61 | 32.68 | 27.35 | 32.90 | 31.40 |
| Qwen3-VL-4B | 0.61 | 15.12 | 9.46 | 4.30 | 16.66 |
| Qwen3.5-9B | 0.36 | 11.49 | 6.42 | 4.03 | 13.41 |
| InternVL2.5-8B | 0.98 | 12.31 | 9.09 | 1.54 | 13.06 |
| LLaVA-NeXT-Mistral-7B | 5.61 | 17.77 | 14.41 | 4.83 | 18.04 |
| Qwen3-VL-30B | 8.86 | 30.78 | 26.81 | 24.18 | 30.62 |
| Gemma3-27B | 4.88 | 26.14 | 21.63 | 19.58 | 27.37 |
Gemini-2.5-Flash leads BLEU, ROUGE-L and CIDEr, while GPT-5.1 leads METEOR and chrF++. chrF++ splits the models into two bands: closed-source and larger open-source models score 27.4–35.6, and smaller open-source models score 13.1–18.0.
Multiple choice: rankings flip
| Model | Direct | + CoT |
|---|---|---|
| GPT-5.1 | 73.65 | 76.80 |
| Gemini-2.5-Flash | 73.20 | 76.05 |
| Qwen3.5-9B | 69.10 | – |
| Qwen3-VL-4B | 68.40 | – |
| GPT-5 mini | 61.90 | 67.00 |
| Qwen3-VL-30B | 60.50 | – |
| Gemma3-27B | 55.65 | – |
| InternVL2.5-8B | 47.55 | – |
| LLaVA-NeXT-Mistral-7B | 27.35 | – |
Overall accuracy (%) on EK-BanglaMCQ, where chance is 25%. Qwen3.5-9B beats GPT-5 mini by 7.20 points here, despite trailing it by 17.44 chrF++ points on captioning. Free-form Bangla generation and grounded answer selection test different skills, so the two tracks are not interchangeable. Chain-of-thought prompting adds 2.85–5.10 points for the closed-source models.
Bangla vs. English on identical clips


- All nine models score lower in Bangla, from −1.20 points for Gemini-2.5-Flash to −23.90 for InternVL2.5-8B.
- Averaged over models, Object questions show the largest gap (10.48 points) and Hand-Object Relation the smallest (5.50 points).
- Without any frames, GPT-5 mini and GPT-5.1 reach only 31.4% and 35.5%, close to the 25% chance level, which confirms that the verifier removed most questions answerable from text alone.
Qualitative examples

Takeaways
Fine-grained multimodal ability does not transfer evenly across languages or evaluation formats. Most caption errors are visual (hallucination, omission, wrong verb, wrong object), not linguistic, so the Bangla gap is about grounding, not just fluency. EK-Bangla gives a diagnostic baseline for measuring progress on low-resource egocentric understanding.