← ICLR 2024 leaderboard

Human-in-the-Loop Test-Time Domain Adaptation for Object Detection

Dzung Anh Doan, Bach Long Nguyen, Terry Lim, Madhuka Jayawardhana, Ian Reid, Markus Wagner, Tat-Jun Chin

transfer & meta learninghuman-in-the-loop machine learningtest-time domain adaptationobject detection
13.70100
Fused
band ≈ ±14 pct pts (from σ = 0.28)
12.50100
Mimo
band ≈ ±19 pct pts (from σ = 0.38)
16.10100
DeepSeek
band ≈ ±21 pct pts (from σ = 0.41)

OpenReview ground truth

Rejected

TL;DR — Introducing human-in-the-loop test-time domain adaptation for object detection

Abstract

Prior to deployment, an object detector is trained on a dataset compiled from a previous data collection campaign. However, the environment in which the object detector is deployed will invariably evolve, particularly in outdoor settings where changes in lighting, weather and seasons will significantly affect the appearance of the scene and target objects. It is almost impossible for all potential scenarios that the object detector may come across to be present in a finite training dataset. This necessitates continuous updates to the object detector to maintain satisfactory performance. Test-time domain adaptation techniques enable machine learning models to self-adapt based on the distributions of the testing data. However, existing methods mainly focus on fully automated adaptation, which make sense for applications such as self-driving cars. Despite the prevalence of full automated approaches, in some applications such as surveillance, there is usually a human operator overseeing the system's operation. We propose to involve the operator in domain adaptation to raise the performance of object detection beyond what is achievable by fully automated adaptation. To reduce manual effort, the proposed method only requires the operator to provide weak labels, which are then used to guide the adaptation process. Furthermore, the proposed method can be performed online, facilitating its applications in scenarios where inference and domain adaptation must be carried out simultaneously. Our experiments show that the proposed method outperforms existing works, demonstrating a great benefit of human-in-the-loop test-time domain adaptation.

Author context

Most prolific author: 4 submissions (credibility 1.00).

No mass-submission penalty for this paper (authors within normal submission volume).

Aggregate statistics only — no individual author rankings.

Ranking trajectory

Percentile by tournament round — convergence indicates rating stability.

Battle history — 34 comparisons

Ranked above opponent in 42% of matchups.

Judge assessments

Mean overall score 0.0 ± 0.0 (n = 34)