Ensuring AI Language Equity in Australian Public Services
- 2026 Global Voices

- 8 minutes ago
- 11 min read
Eve Carlin, University of Sydney, AI For Good Fellow 2026
Executive Summary
Artificial intelligence (AI) is increasingly embedded across Australian Government service delivery, including through chatbots, virtual assistants, and automated decision-support systems. As agencies expand multilingual AI capabilities, evidence indicates system performance remains uneven across languages, particularly in lower-resource and culturally specific contexts. In non-English languages, AI systems may produce fluent but inaccurate information, creating risks in high-stakes areas such as eligibility, compliance, and entitlements.
This paper recommends amending the Australian Government AI Technical Standard (Technical Standard) to introduce explicit multilingual failure conditions under Statement 22, Criterion 80. Agencies would be required to define minimum performance thresholds for each supported language before deployment. Where systems cannot meet these thresholds, automated outputs would be blocked and users redirected to a human or interpreter pathway, such as TIS National.
Implementation would be led by the Digital Transformation Agency (DTA) through updates to the Technical Standard and associated guidance. Commonwealth agencies would integrate language thresholds and fallback pathways into existing and new systems over a staged 12-month period. Costs are expected to be low and primarily administrative, absorbed within existing agency budgets.
Key barriers include the technical difficulty of measuring AI performance consistently across languages and reliance on interpreter service capacity where automated outputs are suppressed. The reform also introduces a trade-off between accessibility and reliability, as some automated services may be restricted in languages where minimum standards cannot be met. By embedding language-specific assurance mechanisms into AI governance, the proposal would help ensure digital public services do not deepen inequities for culturally and linguistically diverse (CALD) communities.
Problem Identification
AI-mediated systems are already used across major Commonwealth service platforms, shaping how people access information, navigate entitlements, and interact with government. Current applications include:
Services Australia AI-powered digital assistants used across Medicare, Centrelink, Child Support, and myGov
Australian Taxation Office virtual assistant 'Alex', handling millions of taxpayer queries annually (Hirschhorn 2021)
Department of Home Affairs (Home Affairs) AI-supported visa triage and document processing systems
Department of Veterans' Affairs chatbot pilots helping veterans navigate entitlements
However, training datasets for these systems often over-represent widely used languages and lack sufficient cultural and contextual nuance, increasing the risk of inaccurate or misleading outputs in non-English contexts (OECD, 2025, p. 41). Unlike traditional translated documents, AI-generated outputs are dynamic and probabilistic. Errors may therefore appear fluent and authoritative, making them difficult for users to identify. These failures create significant risks in high-stakes service environments. Community stakeholders have reported cases where machine-translated government information was unclear, incorrect, or culturally inappropriate (Soldatic et al., 2025). This may result in individuals misunderstanding eligibility criteria, missing legal obligations, or disengaging from essential services altogether.
The issue affects a substantial proportion of the Australian population. More than 5.8 million Australians, or 22.8 per cent of the population, speak a language other than English at home, while approximately 700,000 overseas-born residents report limited English proficiency (ABS, 2021). These communities also report higher rates of digital service use than English-only speakers, meaning CALD communities are among the groups most exposed to failures in AI-mediated systems (Australian Digital Inclusion Index, 2025, p. 19).
If left unaddressed, these failures risk entrenching a two-tiered system in which access to welfare, healthcare, migration support, and taxation services is increasingly shaped by language and cultural background rather than legal entitlement (do Carmo, 2025).
Context
Background
Australia is rapidly embedding AI into public service delivery, with automated systems now guiding people through welfare, health, migration, and taxation processes (Services Australia, 2025). AI-enabled tools can increase the linguistic accessibility of public services; however, multilingual AI systems do not perform evenly across languages, particularly where training data is limited or culturally specific communication is required (Nicholas and Bhatia 2023).
The core issue is not only uneven multilingual AI performance, but the absence of governance mechanisms that define, measure, and enforce acceptable standards across languages. Existing assurance frameworks do not require agencies to establish language-specific performance benchmarks or failure thresholds before deployment. As a result, AI systems are often trained and evaluated using datasets that disproportionately reflect English-language usage. Lower-resource languages frequently lack sufficient training material, evaluation benchmarks, and culturally contextual data. This contributes to weaker and less tested performance, particularly when systems encounter culturally specific, administrative, or legal terminology.
Current Policy Landscape
Australia has established multiple frameworks governing AI assurance, language access, and translation quality across public service delivery (Table 1). However, responsibility for multilingual AI governance remains fragmented across technical, accessibility, and translation frameworks, and no integrated mechanism currently requires agencies to verify multilingual AI performance before deployment.

Current Commonwealth AI governance frameworks contain no explicit requirements for multilingual assurance. Agencies are not required to:
• Test AI systems on a language-by-language basis.
• Establish multilingual performance benchmarks.
• Monitor multilingual performance after deployment.
• Restrict outputs where language-specific standards are not met.
• Redirect users to human or interpreter-supported services when necessary.
Consequently, multilingual capability is assumed rather than demonstrated.
Relevant Precedents
Comparable jurisdictions demonstrate both the risks of inadequate multilingual safeguards and the feasibility of stronger governance measures. In the United Kingdom, the use of AI and machine translation in healthcare has produced documented errors in lower-resource languages, reducing accessibility and increasing the risk of inaccurate information being provided to users (do Carmo, 2025). The Western Australian Government has identified similar quality, accuracy, and cultural alignment issues in AI translation tools, indicating that these challenges are not limited to overseas contexts (Office of Multicultural Interests, 2025). Together, these examples highlight a common problem: multilingual capability is often assumed rather than independently verified.
Canada demonstrates a potential governance response. The Treasury Board of Canada Secretariat's Directive on Automated Decision-Making requires federal agencies to undertake structured pre-deployment risk assessments for automated systems. While not language-specific, it shows that enforceable assurance requirements can be embedded within public sector AI governance and adapted to address multilingual risks (Treasury Board of Canada Secretariat, 2025).
Australia's existing frameworks address AI governance, translation quality, and language accessibility in isolation. Closing this gap will require enforceable multilingual performance thresholds, defined failure conditions, and reliable human fallback pathways.
Policy Options
To ensure equitable access to AI-enabled public services, the Australian Government should require that AI systems perform as safely, accurately, and usefully in supported non-English languages as they do in English before deployment. Where this cannot be achieved, agencies should provide clear human or interpreter pathways.
Option 1: Amend the Technical Standard to include explicit multilingual failure conditions
Criterion 80 of the Technical Standard already requires agencies to identify and address situations when AI outputs should not be provided (DTA, 2024). This option would extend that obligation to language-based performance. Where an AI system cannot demonstrate it meets the quality thresholds established under Criterion 79 for a given language, it must not produce automated outputs in that language and must instead direct the user to a human interpreter service such as TIS National. For generative AI systems, agencies would also be required to implement automatic filters that detect and block outputs in languages where the system's training data is insufficient to meet quality standards.
Estimated costs are low and would be largely absorbed within existing agency resources. The DTA would issue and own the amendment, with individual agencies responsible for resourcing and communicating fallback pathways to users. The key strength of this option is enforceability: placing the multilingual stop rule within an existing Standard removes agency discretion over whether to act. The key risk is that if fallback services such as TIS National are not adequately resourced or clearly surfaced to users at the point of failure, the amendment may reduce automated accessibility without guaranteeing a meaningful alternative.
Option 2: Amend the Multicultural Access and Equity Policy Guide (Policy Guide) to include AI language performance requirements in outsourced service contracts
The Policy Guide already requires agencies to specify multicultural access and equity accountabilities in outsourced contracts, including the provision of appropriate translating and interpreting services (DHA, 2018, p. 11). This option would extend that obligation to AI-mediated service delivery, requiring contracts to include minimum language coverage, measurable performance standards per language, and human quality assurance for AI-generated content. Costs are low and primarily administrative, covering updates to the Guide, revised tender templates, and agency guidance. DHA would own the amendment; individual agencies would be responsible for contractual compliance and annual performance reporting.
The key strength is reach: the policy applies to all non-corporate Commonwealth entities under the PGPA Act and extends to contractors, covering AI language performance across all outsourced government activity without new legislation. The key limitation is that the Guide is principles-based and relies on agency self-reporting, meaning implementation quality may vary without independent audit mechanisms.
Option 3: Amend the Australian Government Language Services Guidelines to extend ISO 18587:2017 post-editing and minimum performance benchmarks to AI-powered public-facing tools
The Guidelines already require agencies deploying machine translation to comply with ISO 18587:2017 and engage NAATI-credentialed translators to post-edit output. The Guidelines also recognise that machine-translated content ‘may be less reliable or not viable for minor languages’ (DHA, 2019, p. 46). However, Section 6 currently applies only to static translated documents and does not extend to AI-powered tools such as virtual assistants and automated chat interfaces. This amendment would extend Section 6 to cover these tools, requiring agencies to define minimum performance benchmarks for every language in which an AI system operates, conduct ongoing quality assurance using NAATI-credentialed translators, and suspend AI outputs in a language where benchmarks are not met.
Estimated costs would involve updates to Section 6, development of benchmark methodology guidance, and ongoing NAATI-credentialed quality assurance at the agency level, largely absorbed within existing departmental resourcing. DHA, as custodian of the Guidelines, would implement the amendment. The key strength is technical specificity: using existing international standards and credentialing systems establishes clear performance criteria. The key limitation is that the Guidelines remain non-binding, meaning effectiveness would depend on agencies voluntarily incorporating the benchmarks into procurement and operational processes.
Policy Recommendation
Option 1, to amend the Australian Government AI Technical Standard to include explicit multilingual failure conditions, is recommended as the most effective approach. Unlike contractual requirements or non-binding guidance, it embeds a mandatory technical safeguard within the existing whole-of-government assurance framework.
The amendment would require agencies to establish language-specific performance thresholds and implement language detection mechanisms to assess whether a user's language falls within the system's validated capability. Where those thresholds are met, the system would operate normally. Where they are not, the system would be required to suppress automated outputs and redirect users to an appropriate human service or interpreter pathway, such as TIS National. For generative AI systems, this would include controls preventing output generation in unsupported or low-confidence languages.
By defining multilingual underperformance as a failure condition under Criterion 80, linked to performance thresholds established under Criterion 79 of the Australian Government AI Technical Standard (DTA, 2024), the amendment would shift assurance from passive monitoring to an active safeguard that prevents unsafe or unreliable outputs from being delivered to users.
Implementation

Costing and Resourcing
The estimated cost of this reform is low to moderate, primarily involving system updates and fallback pathway integration. International comparators indicate that policy expenditure for similar reforms is minimal, with most of the implementation burden falling on agencies and absorbed into the existing budgets.
The closest governance analogue is Canada’s Directive on Automated Decision-Making, which introduced mandatory AI oversight and human intervention requirements through a 12-month rollout without a separate implementation budget (Treasury Board of Canada Secretariat 2025).
For the DTA, expenditure would primarily relate to technical guidance, updates to the Australian Government AI Technical Standard, and implementation support for agencies. As the reform builds on existing assurance and accessibility frameworks rather than creating new governance structures, implementation costs would likely remain low and largely administrative in nature, estimated at approximately $250,000–$750,000 over two years. This is consistent with existing Australian Government AI governance measures, including the 2024–25 ‘Supporting safe and responsible AI’ measure, which the DTA stated would be delivered using existing APS resources (Australian Government, 2024). Agency-level system changes, including language testing and fallback pathway integration, would be absorbed within existing digital service budgets and operational resources.
Success Metrics
Success would be measured through indicators focused on multilingual reliability, compliance, and safe service delivery, including:
100 per cent of in-scope AI systems implementing documented multilingual failure conditions, language-specific thresholds, and fallback pathways within 12 months.
Independent verification that supported languages meet minimum performance standards before deployment, with annual review requirements for all high-impact systems.
Accurate escalation to fallback pathways in cases where multilingual thresholds are not met, with agencies targeting near-complete compliance through periodic auditing.
Year-on-year reductions in complaints or reported harms relating to inaccurate multilingual AI outputs in public service contexts.
Evaluation would occur through agency reporting and cross-government review led by the DTA, with input from DHA.
Barriers and Risks
Barriers
A key barrier to implementation is the technical complexity of defining and enforcing language-specific performance thresholds. Many agencies do not currently measure system performance at a language-disaggregated level, and introducing threshold-based controls requires new evaluation processes and technical infrastructure. This can be addressed through central guidance and shared methodologies published by the DTA, but it nonetheless represents a meaningful shift in how agencies assess AI system performance.
A second barrier is dependency on interpreter and human service capacity. Where automated outputs are suppressed, users must be redirected to alternative pathways. If services such as TIS National are not adequately resourced or integrated into digital workflows, the policy may reduce automated accessibility without guaranteeing a timely or effective alternative. This could be mitigated through phased implementation, stronger integration between digital platforms and human support services, and tiered fallback pathways based on service risk. Low-risk matters could use multilingual guidance or callback systems, while high-risk matters would be prioritised for interpreter escalation.
A third barrier is the technical difficulty of accurately identifying user language, particularly in cases of code-switching or dialect variation. This introduces implementation complexity and requires careful system design and testing before deployment.
Risks
Social risk: Some language groups may experience reduced access to automated services where systems cannot meet performance thresholds. While this prevents harm from incorrect outputs, it may create perceptions of unequal service access if fallback pathways are slower or less convenient than English-language services. Over time, this could reduce trust in digital government services among CaLD communities. This risk could be mitigated through phased implementation, minimum standards for fallback pathways, and clear public communication explaining why automated outputs have been restricted.
Political and economic risk: Restricting AI functionality in certain languages may be perceived as evidence of system underperformance, while greater reliance on interpreters may increase operational costs over time. Agencies may face pressure to narrow language coverage or weaken threshold requirements to reduce costs and maintain perceptions of efficiency. There is also a risk the policy is framed as slowing innovation in government service delivery. These risks could be mitigated through shared multilingual testing infrastructure coordinated by the DTA, targeted funding support for high-demand services, and public framing that positions stop rules as a safety safeguard rather than a limitation on access.
References
Australian Bureau of Statistics (ABS). 2021. Census of Population and Housing. Canberra: Australian Government.
Australian Government. 2024. 2024–25 Budget Estimates briefing pack: Digital Transformation Agency. Department of Finance.
Australian Government Department of Finance. 2024. National Framework for the Assurance of Artificial Intelligence in Government. Canberra: Australian Government.
Australian Government Department of Finance. 2024. Commonwealth Procurement Rules. Canberra: Australian Government.
Australian Human Rights Commission. 2023. The Need for Human Rights-Centred Artificial Intelligence: Submission to the Department of Industry, Science and Resources. Sydney: Australian Government.
Department of Health and Aged Care. 2025. Safe and Responsible Artificial Intelligence in Health Care: Legislation and Regulation Review – Final Report. Canberra: Australian Government.
Department of Home Affairs. 2018. Multicultural Access and Equity Policy Guide. Canberra: Australian Government.
Department of Home Affairs. 2019. Australian Government Language Services Guidelines. Canberra: Australian Government.
Digital Transformation Agency. 2024. Australian Government AI Technical Standard. Canberra: Australian Government.
Digital Transformation Agency. 2024. Policy for the Responsible Use of Artificial Intelligence in Government. Canberra: Australian Government.
do Carmo, Fernanda. 2025. Evidence Review Report on the Use of AI for Multilingual Communication in Public Services: With a Specific Focus on the NHS. Guildford: Centre for Translation Studies, University of Surrey. https://doi.org/10.15126/901630
Hirschhorn, Jeremy. 2021. “Digital Transformation: Australia as a World Leader.” Paper delivered to the 2021 Pearcey Oration and Victorian Entrepreneur Award, 1 September 2021. Canberra: Australian Taxation Office.
National Institute of Standards and Technology (NIST). 2024. Artificial Intelligence Risk Management Framework: Generative AI Profile. Gaithersburg, MD: U.S. Department of Commerce.
Nicholas, Gabriel, and Aliya Bhatia. 2023. “Lost in Translation: Large Language Models in Non-English Content Analysis.” arXiv preprint arXiv:2306.07377. https://doi.org/10.48550/arXiv.2306.07377
Office of Multicultural Interests. 2025. Artificial Intelligence in Language Services. Perth: Government of Western Australia.
Organisation for Economic Co-operation and Development (OECD). 2025. Governing with Artificial Intelligence: Are Governments Ready? Paris: OECD Publishing.
Roy Morgan, RMIT University, Swinburne University of Technology, and Centre for Social Impact. 2025. Australian Digital Inclusion Index 2025. Melbourne: RMIT University.
Services Australia. 2025. Automation and AI Strategy 2025–27. Canberra: Australian Government.
Soldatic, Karen, Meagan Lee, Elvira Tunggal, Andrea Liao, and Liam Magee. 2025. 'Rethinking Digital and AI Inclusion: Participatory and Intersectionality-Informed Methods for Disability and Migrant Justice.' Frontiers in Sociology 10: 1593330. https://doi.org/10.3389/fsoc.2025.1593330
Treasury Board of Canada Secretariat. 2025. Directive on Automated Decision-Making. Ottawa: Government of Canada. https://www.tbs-sct.canada.ca/pol/doc-eng.aspx?id=32592
Whitehead, Lisa, Jelena Talevski, Farhad Fatehi, and Allison Beauchamp. 2023.
'Barriers to and Facilitators of Digital Health among Culturally and Linguistically Diverse Populations: Qualitative Systematic Review.' Journal of Medical Internet Research 25: e42719. https://doi.org/10.2196/42719
The views and opinions expressed by Global Voices Fellows do not necessarily reflect those of the organisation or its staff.
.png)


