It is argued that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Iris Marion Young's theories of oppression and structural injustice.
Abstract
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, and research priorities within AI. Benchmarks standardise the assessment of systems and facilitate the creation of leaderboards that reward state-of-the-art performance with prestige, citations, trust, and institutional influence. As the costs of developing competitive AI systems rise, these rewards increasingly concentrate among powerful, industry-funded labs. This paper situates these concerns within Iris Marion Young's theories of oppression and structural injustice. It argues that current benchmarking practices may perpetuate systematic harms affecting various actors in AI research, aligning with four of Young's"faces of oppression". Benchmarking culture is further framed as a source of structural injustice, as these harms emerge from normalised, individually defensible practices and network effects, even without explicit wrongdoing. By reinforcing existing power structures and narrowing possible research trajectories, benchmarking may in fact prevent the field from advancing in epistemically robust and socially beneficial ways.
The paper’s central argumentative shift is to change the narrative from bias mitigation to bias management—treating bias not as a defect to be corrected but as an ongoing condition to be governed.
Gabriela Arriagada-Bruneau· Science and Engineering Ethi...· 0 citations
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and alignment research has a different mission: to ensure that catastrophic failures never occur, under sparse evidence, adversarial dynamics, and fat-tailed risk. We argue that the two domains differ along two analytically independent axes —
capability profile
(demonstrating the absence of hazardous behaviours versus the presence of positive capabilities) and
risk profile
(bounding worst-case outcomes under fat-tailed uncertainty versus optimising average-case performance) — and that mainstream epistemic practices are inadequate on both. Building on a structured synthesis grounded in a preregistered bibliometric baseline, we identify five cross-cutting gap dimensions in current alignment research, including the near-absence of institutionalised independent verification. To address these gaps we propose ECAISA, an Epistemic Code for AI Safety and Alignment comprising eight principles, a three-level scoring rubric, a four-level disclosure ladder that reconciles transparency with information-hazard and commercial-confidentiality constraints, a tiered applicability scheme, an infohazard adjudication procedure, and seven anti-gaming mechanisms. ECAISA does not certify that any AI system is safe; it constrains how safety-relevant research claims are documented, checked, and relied upon, with
auditability rather than certification
as its governance target. A retrospective rubric audit (κ = 0.79) demonstrates instrument feasibility; a four-stage validation roadmap is proposed.
K. Navaie· Artificial Intelligence Revi...· 0 citations
This paper explores how six major technology companies: Amazon, Microsoft, Apple, NVIDIA, Google, and Meta address the environmental implications of their artificial intelligence (AI) infrastructure through their discourse in different genres of official documents. The research draws on corporate texts, including sustainability reports, executive keynotes, shareholder letters, and regulatory filings, along with independent evidence from non-corporate sources. The research investigates three questions. First, it examines how rhetorical strategies vary by communication genre and target audience. Second, it assesses how temporal deferral functions as a means of legitimation across the dataset. Finally, it evaluates whether discursive confidence enabled by temporal deferral correlates with weaker material emissions performance across the sample. Utilising inductive discourse analysis grounded in sociotechnical imaginaries (Jasanoff & Kim, 2009, 2015); the sociology of expectations (Brown & Michael, 2003; Borup et al., 2006); and Crawford's (2021) materialistic critique of computation, the findings indicate a consistent dual narrative structure: an outward-looking promissory register (performed primarily through keynote addresses and shareholder letters) that represents AI as an active ecological actor, and an inward-looking defensive register (limited to regulatory filings) that re-classifies those same obligations as non binding legal risks. Further comparative analysis of this two-register structure against the 2020-2024 emissions trajectories for each corporation indicates a striking asymmetry. Building on prior use of the term"technowashing" (Perakslis, 2020; Ribeiro & Soromenho-Marques, 2022), this paper refines this concept within climate-AI discourse by analyzing its reliance on uncolonized future times where Big Tech uses speculative promises about a future humans have not reached yet to shape the present debate(Brown & Michael, 2003). It argues that this temporal deferral does not merely exaggerate claims; rather, it allows AI-driven promissory discourse to expand alongside rising rates of material extraction, rather than emissions being reduced by tech solutions.
Asra Zahid· Dialogues in Humanities and...· 0 citations
This paper explores how the integration of artificial intelligence (AI) into military decision-making poses new challenges to human rights focused open-source investigation (OSI). It identifies three key epistemic risks emerging from AI’s role in the military ecosystem. Across these risk areas, AI’s perceived neutrality and rationality create feedback loops of impunity, where human oversight is symbolic or compromised—obscuring both the messy realities of warfare and the potential culpability of military decision-makers. The paper then argues that to continue exposing AI-driven military violence, OSI practices must reorient their methodologies. Rather than treating AI as a discrete, technical tool, OSI must focus more systemically and systematically on the dispersed material infrastructures and political forces underpinning AI militarism. By refocusing investigative practices onto these systems, investigators can better reveal concealed networks of accountability linking state and corporate actors. Such a shift also enables OSI to move beyond narrow legalistic frameworks and engage with structural forms of violence that often evade formal accountability. This could support more expansive, justice-oriented investigations—amplifying marginalised perspectives, mapping complicity, and aligning OSI with broader struggles against systemic harm. Detached from the procedural demands of law, OSI can become a more active, politically mobilising practice, amplifying marginalised perspectives, mapping networks of complicity, and supporting grassroots struggles for justice beyond the courtroom.
Despite substantial excitement around the use of AI in law, little information exists on the performance and associated risks of the domain’s widely marketed tools. Recent work, for instance, has demonstrated the significant potential for “hallucinations”—wherein models make up facts, law, and precedent—leading Chief Justice Roberts to spotlight this risk in his annual report on the judiciary. We argue that there is a need for public AI benchmarking in law. First, relative to other AI application domains, the legal AI ecosystem lacks legibility—there is little information about the design and performance of many commercial legal AI systems. Legal AI has not benefited from the types of benchmarking that have catalyzed, measured, and informed AI innovation and responsible use in other domains. Second, we articulate the challenges of the institutional design of benchmarking. We illustrate how benchmarks can be captured, watered down, and abused. Careful institutional design around the why, who, what, and how of benchmarking will be critical to navigate difficult tradeoffs of transparency, objectivity, expertise, and resources. Third, addressing legal AI’s illegibility requires matching institutional models to available resources and constraints. Rather than advocating for a single “best” approach to benchmarking, we show how benchmarking strategies depend on available resources.
Neel Guha, Andy K. Zhang, Christine Tsang et al.· Proceedings of the National...· 1 citation
We argue that AI systems used in conducting foreign policy tasks - broadly enacting'statecraft'- should be a priority test case for technical AI governance research. In enacting foreign policy, we refer to the formulation and implementation of external objectives by political actors. Statecraft is a high-consequence deployment domain, with extreme downside risks and structural properties that standard evaluation practices handle poorly. These features include partial observability, unbounded action spaces, contested ground truth, and multidimensional objectives. This paper advocates for a literature-grounded research agenda. Our contribution is threefold: (i) a claim about the structural conditions of foreign policy that combine catastrophic tail risk with technical evaluation complexities, (ii) an ECOSYSTEM review that highlights the asymmetric focus on ASSESSMENT features over ACCESS, VERIFICATION, SECURITY, and OPERATIONALIZATION, and (iii) a demand-side evaluation framework that decomposes foreign-policy workflows into bounded, evaluable sub-tasks with human recombination. As AI systems are already being deployed in the conduct of war and peace, amid limited public evaluation infrastructure from the technical AI governance community, this agenda is an urgent priority.