Signal Bias Probes: A Framework for Active Multi-Group Fairness Auditing Without Model Reconstruction
Summary
This paper addresses the gap between fairness-aware machine learning training, which in practice offers only limited improvement over standard empirical risk minimization, and the resulting need for reliable post-hoc auditing of deployed models. Existing black-box auditing approaches either require reconstructing the model, exposing it to extraction attacks, or directly estimate fairness metrics while providing little insight into which regions of the data distribution drive bias. The authors note that property-specific auditing, which extracts only targeted fairness information without full model reconstruction, has remained poorly understood until now. They introduce a 'bias probe' framework enabling targeted, adaptive queries that reveal bias structure while preserving model confidentiality, built on by an active auditor called ALeBi that efficiently estimates multi-group fairness metrics. The work establishes new sample complexity guarantees governed by a property-specific complexity measure and extends the analysis to adversarial settings where a model owner might strategically obscure bias, demonstrating a fundamental trade-off between confidentiality and reliable auditing.
Classification
Evidence 1
- Efficient Active Auditing of Multi-Group Fairness with Bias Probes arXiv (cs.LG / cs.AI / cs.CY) 2026-09-30 accessed 2026-10-01T23:00:11+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-6059579203ff
