Preprint 2026 Under review Egocentric VideoHand-Object InteractionBenchmarksLow-Resource Languages

EK-Bangla: When Hand-Object Interaction Meets Low-Resource Generation

Rabeya Akter, Md Salman Shamil, A. K. M. Fazley Rabbi, Arpa Deb, Ridwain Islam, Md Adnan Arefeen

Under review at WACV 2027 (Round 2)

Paper & code coming soon

TL;DR

Same video, same model, different language: multimodal LLMs that describe hand-object interactions well in English drop hand, object, and verb details in Bangla. EK-Bangla measures that gap with a caption track and a multiple-choice track built from the same egocentric clips.

  • 9,600human-verified Bangla captions in EK-BanglaCAP
  • 2,000four-way MCQs in EK-BanglaMCQ, from 1,346 verified captions
  • 54.5%of MLLM draft captions needed native-speaker revision
  • up to −23.9pt accuracy drop for Bangla vs. English MCQs
Overview figure for EK-Bangla: When Hand-Object Interaction Meets Low-Resource Generation
Same video, same model, different target language. Left: two clips described by Gemini 3.1 Pro, with errors marked against the human-verified reference. The Bangla (BN) output carries errors the English (EN) output does not. Right: mean MCQ accuracy of nine models across six categories. Bangla trails English in every category, by 5.5 to 10.5 points.

Abstract

Describing hand-object interactions in egocentric video requires resolving the acting hand, manipulated object, contact, and temporal progression details important for user-facing visual systems. Although multimodal large language models (MLLMs) are increasingly evaluated on fine-grained egocentric understanding, their ability to preserve such grounding in low-resource languages remains unclear. We introduce EK-Bangla, a two-track Bangla benchmark built on the EPIC-KITCHENS-100 validation split. EK-BanglaCAP contains 9,600 MLLM-generated Bangla captions verified by native speakers under a nine-category error taxonomy, with 54.5% requiring revision. EK-BanglaMCQ contains 2,000 four-way questions constructed from verified captions through a proposer, a distractor, and a verifier pipeline that rejects 76.07% of candidates, followed by human validation for quality assurance. Across nine open- and closed-source MLLMs, performance varies substantially across captioning and multiple-choice questions: Qwen3.5-9B, for example, trails GPT-5 mini by 17.44% in chrF++ on captioning yet exceeds it by 7.20% accuracy on MCQ. Under matched visual evidence, all nine models perform worse on Bangla than English MCQs, with drops up to 23.90%. These results show that fine-grained multimodal capability does not transfer uniformly across language and evaluation format, motivating diagnostic evaluation of low-resource egocentric understanding.

Fine-grained, but only in English

Motivation

A user-facing visual assistant has to say which hand did what to which object, and in what order, and it has to say it in the user's own language. Egocentric video makes this hard: the evidence often occupies a small part of the frame, hands and objects occlude each other, the camera moves, and the meaning of an action depends on how it unfolds over time.

Fine-grained hand-object benchmarks such as EgoHOIBench and HanDyVQA are English-only, while multilingual and Bangla benchmarks focus on images or coarse scene-level video. Nothing tests whether an MLLM can convey subtle interaction details precisely in a low-resource language. EK-Bangla fills that gap for Bangla.

Two tracks, one set of clips

Benchmark construction

EK-Bangla data-generation pipeline for captions and multiple-choice questions
Data-generation pipeline. (a) EK-BanglaCAP: eight frames plus narration, verb and noun annotations produce a draft Bangla caption, which native speakers verify and correct. (b) EK-BanglaMCQ: verified captions feed a three-stage proposer, distractor and verifier pipeline, followed by human validation.

EK-BanglaCAP: free-form captions

  1. DraftEight frames sampled at full 1920×1080 resolution, together with the EPIC narration, verb and noun as weak context, go to Gemini 3.1 Pro. A first pass records hand roles, contact, motion and action phases. A second pass rewrites this into one fluent Bangla sentence.
  2. VerifyNative Bangla speakers watch every source clip and revise the caption where needed. 5,234 of 9,600 drafts (54.5%) needed correction.
  3. Label errorsEach corrected draft is tagged with a nine-category taxonomy: six visual errors (wrong verb, object, hand or tool, hallucination, omission) and three linguistic ones (unnatural wording, untranslated text, grammar or spelling).

Verified captions average 24.7 words, against 2.9 words for the original EPIC narrations.

EK-BanglaMCQ: multiple choice

  1. ProposeWorking from caption text alone, so every answer traces back to human-verified content, a proposer picks the applicable categories and writes one question per category.
  2. DistractA distractor generator writes three plausible but wrong options that match the correct answer in form, detail and length.
  3. VerifyA verifier sees only the question and shuffled options, with no video. Any item it answers correctly is rejected, so questions solvable from priors or lexical cues are filtered out. 6,358 of 8,358 proposals (76.07%) were rejected.

A 300-item human audit (15% of the benchmark), with each item judged by two of three annotators, approved 100% of questions as valid, 100% of answers as correct, and 98% of distractors as plausible.

CategoryWhat it testsQuestions
ActionManipulation performed by the hands332
ObjectObject or tool being acted on780
Hand-Object RelationWhich hand contacts which object206
TemporalOrder in which events occur204
LocationSource or destination of the action454
StateObservable change in an object24
Caption length histogram, caption error counts, and MCQ category distribution
Benchmark statistics. (a) Reference caption lengths. (b) Errors found during human verification. Hallucination, omission, wrong verb and wrong object far outnumber linguistic errors. (c) MCQ category distribution.

Results

Nine MLLMs · 8 frames at 456×256

We evaluate three closed-source models (GPT-5.1, GPT-5 mini, Gemini-2.5-Flash) and six open-source models from 4B to 30B parameters.

Captioning: no single leader

ModelBLEU-4ROUGE-LMETEORCIDErchrF++
GPT-5.15.7728.0230.9012.7535.58
GPT-5 mini3.3423.2523.649.4630.85
Gemini-2.5-Flash9.6132.6827.3532.9031.40
Qwen3-VL-4B0.6115.129.464.3016.66
Qwen3.5-9B0.3611.496.424.0313.41
InternVL2.5-8B0.9812.319.091.5413.06
LLaVA-NeXT-Mistral-7B5.6117.7714.414.8318.04
Qwen3-VL-30B8.8630.7826.8124.1830.62
Gemma3-27B4.8826.1421.6319.5827.37

Gemini-2.5-Flash leads BLEU, ROUGE-L and CIDEr, while GPT-5.1 leads METEOR and chrF++. chrF++ splits the models into two bands: closed-source and larger open-source models score 27.4–35.6, and smaller open-source models score 13.1–18.0.

Multiple choice: rankings flip

ModelDirect+ CoT
GPT-5.173.6576.80
Gemini-2.5-Flash73.2076.05
Qwen3.5-9B69.10–
Qwen3-VL-4B68.40–
GPT-5 mini61.9067.00
Qwen3-VL-30B60.50–
Gemma3-27B55.65–
InternVL2.5-8B47.55–
LLaVA-NeXT-Mistral-7B27.35–

Overall accuracy (%) on EK-BanglaMCQ, where chance is 25%. Qwen3.5-9B beats GPT-5 mini by 7.20 points here, despite trailing it by 17.44 chrF++ points on captioning. Free-form Bangla generation and grounded answer selection test different skills, so the two tracks are not interchangeable. Chain-of-thought prompting adds 2.85–5.10 points for the closed-source models.

Bangla vs. English on identical clips

Heatmap of Bangla minus English accuracy per model and category
Bangla minus English accuracy. Frames, prompts and decoding are held fixed. Only the question language changes.
Bar chart of MCQ accuracy as the number of frames grows from 0 to 16
Frame budget. The first frame gives the largest gain. GPT-5 mini peaks at 8 frames, while GPT-5.1 keeps improving up to 16.
  • All nine models score lower in Bangla, from −1.20 points for Gemini-2.5-Flash to −23.90 for InternVL2.5-8B.
  • Averaged over models, Object questions show the largest gap (10.48 points) and Hand-Object Relation the smallest (5.50 points).
  • Without any frames, GPT-5 mini and GPT-5.1 reach only 31.4% and 35.5%, close to the 25% chance level, which confirms that the verifier removed most questions answerable from text alone.

Qualitative examples

Three EK-BanglaMCQ examples with model predictions
EK-BanglaMCQ samples. (a) Glasses held in both hands are clearly visible, and all nine models answer correctly. (b) The evidence occupies a small region, and only four of nine models are correct. (c) Only GPT-5.1 is correct, and chain-of-thought fixes Gemini-2.5-Flash's answer.

Takeaways

Fine-grained multimodal ability does not transfer evenly across languages or evaluation formats. Most caption errors are visual (hallucination, omission, wrong verb, wrong object), not linguistic, so the Bangla gap is about grounding, not just fluency. EK-Bangla gives a diagnostic baseline for measuring progress on low-resource egocentric understanding.