arXiv 2025 Under review Computer VisionAmodal CountingMultimodal

Counting Through Occlusion: Framework for Open World Amodal Counting

Safaeid Hossain Arib1, Rabeya Akter1, Abdul Monaf Chowdhury1, Md Jubair Ahmed Sourov1, Md Mehedi Hasan1

1Department of Robotics and Mechatronics Engineering, University of Dhaka

arXiv preprint arXiv:2511.12702 · Submitted to WACV 2027

TL;DR

Under occlusion, a backbone encodes the occluder, not the objects behind it. CountOCC rebuilds features at occluded locations from visible fragments plus text and exemplar priors, and trains the occluded view to attend like the clean view.

  • −20.8%test MAE on FSC-147-OCC vs. CountGD (−26.7% on validation)
  • −49.9%MAE on CARPK-OCC (9.28 → 4.65)
  • −28.8%MAE on CAPTURe-Real (14.97 → 10.66)
Overview figure for Counting Through Occlusion: Framework for Open World Amodal Counting
Counting people through occlusion. CountGD (left) misses players hidden behind others; CountOCC (right) recovers them.

Abstract

Object counting has achieved remarkable success on visible instances, yet state-of-the-art (SOTA) methods fail under occlusion. This failure stems from a fundamental architectural limitation where backbone networks encode occluding surfaces rather than target objects, thereby corrupting the feature representations required for accurate enumeration. To address this, we present CountOCC, an amodal counting framework that explicitly reconstructs occluded object features through hierarchical multimodal guidance. Rather than accepting degraded encodings, we synthesize complete representations by integrating spatial context from visible fragments with semantic priors from text and visual embeddings, generating features at occluded locations across multiple pyramid levels. We further introduce a visual equivalence objective that enforces consistency in attention space, ensuring that both occluded and unoccluded views of the same scene produce spatially aligned gradient-based attention maps. Together, these complementary mechanisms preserve discriminative properties essential for accurate counting under occlusion. For rigorous evaluation, we establish occlusion-augmented versions of FSC-147 and CARPK (FSC-147-OCC and CARPK-OCC). CountOCC achieves SOTA performance on FSC-147-OCC with 26.72% and 20.80% MAE reduction over prior baselines under occlusion in validation and test, respectively. CountOCC also demonstrates exceptional generalization by setting new SOTA results on CARPK-OCC with 49.89% MAE reduction and on CAPTURe-Real with 28.79% MAE reduction, validating robust amodal counting.

The occlusion problem

Motivation

Four panels: a tray of 12 donuts, the same tray with two donuts occluded, a prior method counting only 10, and CountOCC counting all 12
Amodal counting. With two of twelve donuts hidden, prior methods count only the visible ten. CountOCC predicts the two occluded instances as well.

Method

Reconstruct, then align

CountOCC architecture with image and text encoders, Feature Reconstruction Module, VisEQ, feature enhancer and cross-modality decoder
CountOCC architecture. The Feature Reconstruction Module (FRM) replaces corrupted occluded tokens at each pyramid level; VisEQ aligns gradient-based attention maps of the occluded (student) and original (teacher) views. Reconstructed features flow through the feature enhancer and cross-modality decoder to produce visible and occluded counts.
  1. Feature Reconstruction ModuleLearnable queries at occluded positions self-attend, cross-attend to visible tokens for spatial context, then cross-attend to fused text–exemplar embeddings for semantics, producing class-discriminative features where the occluder was.
  2. Visual Equivalence (VisEQ)A teacher sees the clean image, a student sees the occluded one. Attention-similarity and ROI-consistency losses make their gradient-based attention maps agree, so localization does not depend on occlusion.
  3. New benchmarksFSC-147-OCC and CARPK-OCC add controlled occlusion to standard open-world and car-counting datasets for rigorous amodal evaluation.

Results

FSC-147-OCC · MAE / RMSE · lower is better

MethodPromptVal MAEVal RMSETest MAETest RMSE
CLIP-CountText26.3180.4523.90108.57
CounTXText24.8175.5823.04113.83
CounTRExemplars23.1466.7822.25104.75
LOCAExemplars17.1344.2516.7778.41
CountGDExemplars + Text15.8354.3814.4285.40
CountOCCExemplars + Text11.6035.4011.4238.68

Generalization

BenchmarkCountGD MAECountOCC MAEReduction
CARPK-OCC (test)9.284.65−49.9%
CAPTURe-Real14.9710.66−28.8%
Qualitative comparison of CLIP-Count, CounTX, CounTR, LOCA, CountGD and CountOCC on occluded FSC-147 images
Qualitative comparison on FSC-147-OCC. Columns show CLIP-Count, CounTX, CounTR, LOCA, CountGD, and CountOCC. Prior methods undercount hidden objects; CountOCC counts correctly under occlusion across diverse categories.

Citation

@article{arib2025countocc,
  title   = {Counting Through Occlusion: Framework for Open World Amodal Counting},
  author  = {Arib, Safaeid Hossain and Akter, Rabeya and Chowdhury, Abdul Monaf and
             Sourov, Md Jubair Ahmed and Hasan, Md Mehedi},
  journal = {arXiv preprint arXiv:2511.12702},
  year    = {2025}
}