Project thumbnail and framework overview.
The offline stage mines error slices and freezes slice-to-action mappings. The deployment stage performs ground-truth-free risk screening, Top-2 slice retrieval, unique repair-action execution, and prediction fusion.
Representative post-training repair probes.
The six panels show the annotated scene, baseline output, class-wise confidence thresholding, contextual rescoring, high-resolution inference, and overlapping tiled inference. Score-level probes modify existing predictions, whereas inference-level probes can introduce new detection geometry. This qualitative figure illustrates the repair mechanisms rather than reporting aggregate measurements.
Top-2 complementary repair and prediction fusion.
High-resolution and tiled inference recover different object candidates. The fusion stage preserves useful baseline evidence, combines complementary detections, suppresses duplicates, and refines box coordinates. This figure is a qualitative mechanism illustration.
|
Counterfactual Error-Slice Diagnosis for Improving Object Detection
ABSTRACT:
Average precision is the standard measure of overall object-detection quality, but it does not describe behavior across diagnostic subgroups or determine which post-training action is appropriate for a given image cohort. We present a fixed-detector diagnosis-to-repair framework that complements dataset-level evaluation with error-slice analysis.
The framework organizes detector behavior into connected diagnostic tables, identifies interpretable and sufficiently supported error slices, and measures controlled post-training interventions on their associated complete-image cohorts. These cohort-level responses establish a frozen slice-to-action profile without updating detector weights; they do not imply that every row-level failure represented by a slice is repaired. At inference time, ground-truth-free camera-channel metadata and baseline-prediction features support fixed-budget risk screening and Top-2 error-slice retrieval. The unique actions associated with the retrieved slices are activated, and their outputs are combined with baseline predictions through prediction fusion.
On the designated 9,752-image nuImages test split, the fixed-detector baseline achieves AP50:95 = 0.37279. Under a 10% selected-image budget, Top-1 baseline-plus-action fusion reaches 0.38872 and nonredundant-slice Top-2 prediction fusion reaches 0.39884 at a runtime-normalized cost of 1.464. Extending retrieval to Top-3 produces a slightly lower point estimate of 0.39855 and increases mean latency from 31.21 to 35.80 ms per image. Top-2 therefore has the highest AP50:95 point estimate among the evaluated retrieval depths and lower latency than Top-3, while average precision remains the primary detection measure. Selected images are forcibly routed to eligible non-baseline slices, so the evaluated policy does not provide distance-based abstention or guaranteed image-level diagnosis.
SUMMARY (中文總結):
平均精確率是衡量物件偵測整體品質的標準指標,但無法描述不同診斷子群中的模型行為,也無法判斷特定影像群組適合採用何種訓練後介入方法。本研究提出一套固定偵測器的診斷至修復框架,以錯誤切片分析補充資料集層級的整體評估。
此框架將偵測器行為整理為相互連結的診斷資料表,辨識可解釋且具足夠支持度的錯誤切片,並在其對應的完整影像群組上量測受控訓練後介入方法。這些群組層級反應用於建立固定的切片至動作設定,但不代表切片所涵蓋的每一筆錯誤均已修復。推論階段使用可取得的相機通道資訊與基準預測特徵進行固定預算風險排序及前二名錯誤切片檢索,再啟用檢索切片所對應的不重複動作;若兩個切片對應相同動作,該動作僅執行一次,最後透過預測融合將其輸出與基準預測結合。
在指定的 9,752 張 nuImages 測試影像上,固定偵測器基準的 AP50:95 為 0.37279。在選取 10% 影像的預算下,前一名基準加動作融合達到 0.38872,非冗餘切片前二名預測融合達到 0.39884,正規化執行成本為 1.464。將檢索深度增加至前三名後,點估計略降至 0.39855,平均延遲亦由每張影像 31.21 ms 增加至 35.80 ms;因此,前二名在本研究所評估的檢索深度中取得最高的 AP50:95 點估計,且延遲低於前三名。通過風險篩選的影像會被強制路由至可執行的非基準切片,因此本研究評估的政策不提供基於距離的棄權機制,也不保證影像層級的精確診斷。
RESULTS:
All selective policies use the same fixed 10% routing budget. AP50:95 is the primary detection measure; runtime cost is normalized by the baseline wall-clock time.
| Policy |
AP50:95 |
Delta AP |
APsmall |
Runtime Cost |
FPS |
| Baseline 640 | 0.37279 | 0.00000 | 0.11466 | 1.000 | 46.905 |
| All-image 960 | 0.39832 | 0.02553 | 0.17895 | 7.357 | 6.376 |
| Top-1 replacement | 0.38571 | 0.01292 | 0.13520 | 1.338 | 35.063 |
| Top-1 fusion | 0.38872 | 0.01593 | 0.14580 | 1.362 | 34.447 |
| Top-2 prediction fusion | 0.39884 | 0.02605 | 0.16570 | 1.464 | 32.041 |
| Top-3 prediction fusion | 0.39855 | 0.02576 | 0.16520 | 1.679 | 27.933 |
- Top-2 provides the highest AP50:95 point estimate among the evaluated ground-truth-free retrieval depths.
- High-resolution inference is the strongest individual repair probe, while targeted Top-2 routing approaches its AP with substantially lower runtime cost.
- Top-3 adds computation but slightly reduces the AP point estimate, indicating that additional retrieved slices may introduce redundant or conflicting evidence.
- The detector weights remain unchanged throughout all post-training diagnosis and repair experiments.
REFERENCES:
-
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick,
"Microsoft COCO: Common Objects in Context," in Computer Vision - ECCV 2014, pp. 740-755, 2014.
doi:10.1007/978-3-319-10602-1_48
-
D. Bolya, S. Foley, J. Hays, and J. Hoffman,
"TIDE: A General Toolbox for Identifying Object Detection Errors," in Computer Vision - ECCV 2020, pp. 558-573, 2020.
doi:10.1007/978-3-030-58580-8_33
-
Y. Chung, T. Kraska, N. Polyzotis, K. H. Tae, and S. E. Whang,
"Slice Finder: Automated Data Slicing for Model Validation," in 2019 IEEE 35th International Conference on Data Engineering, pp. 1550-1553, 2019.
doi:10.1109/ICDE.2019.00139
-
G. Jocher and J. Qiu, "Ultralytics YOLO11," software, version 11.0.0, 2024.
Official repository
-
Motional, "nuImages," official dataset website, 2020.
Dataset website
-
J. H. Friedman, "Greedy Function Approximation: A Gradient Boosting Machine," The Annals of Statistics, vol. 29, no. 5, pp. 1189-1232, 2001.
doi:10.1214/aos/1013203451
-
P. C. Mahalanobis, "On the Generalised Distance in Statistics," Proceedings of the National Institute of Sciences of India, vol. 2, no. 1, pp. 49-55, 1936.
-
N. Bodla, B. Singh, R. Chellappa, and L. S. Davis,
"Soft-NMS - Improving Object Detection with One Line of Code," in 2017 IEEE International Conference on Computer Vision, pp. 5561-5569, 2017.
doi:10.1109/ICCV.2017.593
|