arXiv 2025 Under review Computer VisionAmodal CountingMultimodal

Counting Through Occlusion: Framework for Open World Amodal Counting

Safaeid Hossain Arib1, Rabeya Akter1, Abdul Monaf Chowdhury1, Md Jubair Ahmed Sourov1, Md Mehedi Hasan1

1Department of Robotics and Mechatronics Engineering, University of Dhaka

arXiv preprint arXiv:2511.12702 · Submitted to WACV 2027

TL;DR

Under occlusion, a backbone encodes the occluder, not the objects behind it. CountOCC rebuilds features at occluded locations from visible fragments plus text and exemplar priors, and trains the occluded view to attend like the clean view.

  • −20.8%test MAE on FSC-147-OCC vs. CountGD (−26.7% on validation)
  • −68%occluded-region MAE on FSC-147-OCC test (18.16 → 5.80)
  • −49.9%MAE on CARPK-OCC (9.28 → 4.65)
  • −28.8%MAE on CAPTURe-Real (14.97 → 10.66)
Overview figure for Counting Through Occlusion: Framework for Open World Amodal Counting
Amodal counting. With two of twelve donuts hidden, prior methods count only the visible ten. CountOCC also predicts the two occluded instances.

Abstract

Object counting has achieved remarkable success on visible instances, yet state-of-the-art (SOTA) methods fail under occlusion. This failure stems from a fundamental architectural limitation where backbone networks encode occluding surfaces rather than target objects, thereby corrupting the feature representations required for accurate enumeration. To address this, we present CountOCC, an amodal counting framework that explicitly reconstructs occluded object features through hierarchical multimodal guidance. Rather than accepting degraded encodings, we synthesize complete representations by integrating spatial context from visible fragments with semantic priors from text and visual embeddings, generating features at occluded locations across multiple pyramid levels. We further introduce a visual equivalence objective that enforces consistency in attention space, ensuring that both occluded and unoccluded views of the same scene produce spatially aligned gradient-based attention maps. Together, these complementary mechanisms preserve discriminative properties essential for accurate counting under occlusion. For rigorous evaluation, we establish occlusion-augmented versions of FSC-147 and CARPK (FSC-147-OCC and CARPK-OCC). CountOCC achieves SOTA performance on FSC-147-OCC with 26.72% and 20.80% MAE reduction over prior baselines under occlusion in validation and test, respectively. CountOCC also demonstrates exceptional generalization by setting new SOTA results on CARPK-OCC with 49.89% MAE reduction and on CAPTURe-Real with 28.79% MAE reduction, validating robust amodal counting.

Counters only count what they see

Motivation

Open-world counters such as CountGD count any category from a text prompt or a few exemplar boxes. When objects are partly hidden, they fail in a predictable way. The backbone encodes the occluding surface instead of the objects behind it, so the features needed to count those objects are simply missing. Asking a model to count harder does not help, because the evidence is already corrupted.

CountOCC does not accept the degraded features. It reconstructs them, using what is visible around the occluder together with what the text and exemplars say the object should look like. It then checks that the occluded view attends to the same places as a clean view of the scene.

Method

Reconstruct, then align

CountOCC architecture with image and text encoders, Feature Reconstruction Module, VisEQ, feature enhancer and cross-modality decoder
CountOCC architecture. FRM replaces corrupted occluded tokens at every pyramid level. VisEQ aligns gradient-based attention maps of the occluded (student) and original (teacher) views. Reconstructed features flow through the feature enhancer and cross-modality decoder to produce visible and occluded counts.

1. Feature Reconstruction Module (FRM)

At each of three Swin Transformer pyramid levels (256, 512 and 1024 channels), visible tokens are kept and occluded positions are replaced by a learnable mask embedding. These queries self-attend, cross-attend to the visible tokens for spatial context, and then cross-attend to the fused text–exemplar embedding for semantics.

\[ \mathbf{Q}_{\mathrm{vis}} = \Psi_{ca}\big(\Psi_{sa}(\mathbf{Q}_0),\, \mathbf{Z}_{\mathrm{vis}}\big) + \Psi_{sa}(\mathbf{Q}_0), \qquad \mathbf{Z}_{\mathrm{cond}} = \Psi_{ca}(\mathbf{Q}_{\mathrm{vis}},\, \mathbf{Z}_{v,t}) + \mathbf{Q}_{\mathrm{vis}} \]
\[ \hat{\mathbf{Z}}_{\mathrm{occ}} = \Phi_{\mathrm{mlp}}(\mathbf{Z}_{\mathrm{cond}}) + \mathbf{Z}_{\mathrm{cond}}, \qquad \tilde{\mathbf{Z}} = \begin{cases} \mathbf{Z} & \text{visible} \\ \hat{\mathbf{Z}}_{\mathrm{occ}} & \text{occluded} \end{cases} \]

A frozen teacher sees the clean image and provides target features at the occluded positions. Reconstruction is supervised with a Charbonnier, cosine and \(\ell_2\) loss, summed over levels.

\[ \mathcal{L}_{\mathrm{rec}} = \sum_{\ell} \sum_{i \in \mathcal{O}^{(\ell)}} \Big[ \lambda_{\mathrm{charb}} \sqrt{\lVert \Delta^{(\ell)} \rVert_2^2 + \epsilon^2} + \lambda_{\cos}\big(1 - \cos(\hat{\mathbf{Z}}_S^{(\ell)}, \hat{\mathbf{Z}}_T^{(\ell)})\big) + \lambda_{\ell_2} \lVert \Delta^{(\ell)} \rVert_2^2 \Big] \]

2. Visual Equivalence (VisEQ)

FRM fixes the features. VisEQ fixes where the model looks. It computes language-conditioned Grad-CAM maps across pyramid levels for a teacher on the clean image and a student on the occluded one, and aligns them.

\[ \mathcal{L}_{\mathrm{sim}} = \sum_{H,W} \Big[ \lambda_{\ell_2} \lVert \mathbf{G}_T - \mathbf{G}_S \rVert_2^2 + \lambda_{\cos}\big(1 - \cos(\mathbf{G}_T, \mathbf{G}_S)\big) \Big] \]

A region-of-interest consistency loss \(\mathcal{L}_{\mathrm{cst}}\) rewards high, low-variance activation wherever either network is confident, which prevents the trivial solution of both maps going flat.

Feature Reconstruction Module diagram
FRM. Occluded queries gather spatial context from visible tokens, then semantic guidance from text and exemplars.
Visual Equivalence framework diagram
VisEQ. Teacher and student attention maps are aligned across clean and occluded views.

3. New occlusion benchmarks

FSC-147-OCC and CARPK-OCC add controlled occlusion to the FSC-147 open-world counting benchmark (147 categories) and the CARPK car-counting dataset, with separate ground truth for visible and occluded objects.

Sample images from FSC-147-OCC and CARPK-OCC
Benchmark samples from (a) FSC-147-OCC and (b) CARPK-OCC.

Results

MAE / RMSE · lower is better

FSC-147-OCC

MethodPromptVal MAEVal RMSETest MAETest RMSE
CLIP-CountText26.3180.4523.90108.57
CounTXText24.8175.5823.04113.83
CounTRExemplars23.1466.7822.25104.75
LOCAExemplars17.1344.2516.7778.41
CountGDExemplars + Text15.8354.3814.4285.40
CountOCCExemplars + Text11.6035.4011.4238.68

Where the gain comes from

Splitting the error into visible and occluded regions shows CountOCC keeps visible-region accuracy about the same while cutting occluded-region error by roughly a factor of three. The model is reasoning about what is hidden, not just exploiting the occluder's appearance.

MethodVal visible MAEVal occluded MAETest visible MAETest occluded MAE
CountGD8.0517.469.7418.16
CountOCC8.045.618.525.80

Generalization

BenchmarkCountGD MAECountOCC MAECountGD RMSECountOCC RMSE
CARPK-OCC (test)9.284.6511.275.91
CAPTURe-Real14.9710.6641.6241.31
CrowdHuman9.978.2424.4617.87

On the original, unoccluded FSC-147, CountOCC stays competitive with a test MAE of 7.02, second only to CountGD (5.74), so robustness to occlusion does not come at the expense of normal counting.

CountGD and CountOCC density responses on a CrowdHuman basketball scene
CrowdHuman. CountOCC assigns stronger responses to partially occluded people, such as the leftmost player, reducing undercounting in overlapping regions.
Qualitative comparison of CLIP-Count, CounTX, CounTR, LOCA, CountGD and CountOCC on occluded FSC-147 images
Qualitative comparison on FSC-147-OCC. Columns show CLIP-Count, CounTX, CounTR, LOCA, CountGD and CountOCC. Prior methods undercount hidden objects, while CountOCC counts correctly across diverse scenes.

What matters

Ablations on FSC-147-OCC

VariantVal MAEVal RMSETest MAETest RMSE
Design
No FRM (CountGD)15.8354.3814.4285.40
FRM at one level13.1654.5113.77108.63
FRM at all levels11.3248.1211.9091.45
FRM at all levels + VisEQ11.6035.4011.4238.68
Reconstruction loss
\(\ell_2\)13.8878.6713.2488.93
+ cosine12.1848.8812.3887.04
+ Charbonnier11.3248.1211.9091.45
+ VisEQ losses11.6035.4011.4238.68
  • Reconstruct at every scale. FRM at one level helps MAE but not large errors. All three levels cut validation MAE by 28.5%.
  • VisEQ tames the worst cases. Adding attention alignment more than halves test RMSE (91.45 → 38.68), meaning far fewer large miscounts.
t-SNE of occluded, ground-truth and reconstructed features at three pyramid levels
Reconstructed features land where they should. t-SNE at three pyramid levels comparing occluded features (red), ground-truth features from unoccluded images (green) and reconstructed features (blue).

Limitations

Examples where total count is correct but predicted locations in the occluded region differ from true locations
Right count, approximate location. CountOCC predicts the correct totals, but the spatial layout of density inside the occluded region does not always match the true object positions.

FRM recovers features that are informative for how many objects are hidden, but it does not enforce a one-to-one match with where they are, so CountOCC targets amodal counting rather than amodal detection. It also assumes an occlusion mask, which in practice could come from a segmentation model. Without a mask it behaves as a standard open-world counter. Predicting the mask jointly with the count is an important next step.

Citation

@article{arib2025countocc,
  title   = {Counting Through Occlusion: Framework for Open World Amodal Counting},
  author  = {Arib, Safaeid Hossain and Akter, Rabeya and Chowdhury, Abdul Monaf and
             Sourov, Md Jubair Ahmed and Hasan, Md Mehedi},
  journal = {arXiv preprint arXiv:2511.12702},
  year    = {2025}
}