Three AI agents, two countries, and one uneven world wide web
In this news:
Key takeaways
- Three AI agents were tested on updating data for the U.S. and Iran.
- Agents showed stark differences in permission requests and account registration.
- Multilingual performance in Farsi revealed significant information retrieval hurdles.
I’m a technology and human rights researcher. For the past few years, one of my focuses has been the question of how language shapes the ways we benefit from, or are harmed by, AI. I developed an open-source platform for language-pair analysis of LLM responses across different languages and contexts. I’ve also worked on evaluating policy-prompts guardrails and on whether giving LLM guardrails access to tools can make them more reliable and trustworthy (this work was recently accepted to NeurIPS! Yay!!).
Recently, though, I was on a panel at RightsCon on the human rights impact assessment of agentic AI. It got me thinking more about which aspects of language matter when evaluating LLM agents. I wanted to move beyond asking whether a model performs differently when I ask the same question in English versus Farsi (my native language), and look instead at the whole agentic trajectory (reasoning, planning, search, source selection, source hierarchy, artifact creation) while taking language and context into account.
So I decided to run a test.
Human in the loop from repeated permission prompts to almost no intervention
For those of us working in digital rights, human-in-the-loop (HITL) oversight of AI agents joins a longer line of debates about “informed consent,” from GDPR consent requirements to cookie pop-ups and the routine ticking of Terms of Service and Privacy Policy boxes. With that in mind, I paid close attention to how each agent involved me during the experiment.
For accessing and fetching information from websites, GPT asked for permission only once, at the very beginning of the task. It requested access to websites and offered an “allow all relevant sites” option, which I selected. After that, it did not ask again.
Claude asked questions throughout the task, both about accessing websites for research and about extracting material from the World Bank site. Unlike GPT, it offered no “allow all relevant sites” option, so it requested permission each time: nine times for the US task (all .gov websites) and nine times for the Iran task (mainly .ir domains and also fa.wikipedia sources). I approved every request.
Claude also ran into a technical limit that triggered a different kind of HITL moment. For both Iran and the US, it couldn’t load the live GPPD country profiles, because the portal builds its pages with JavaScript, and Claude’s sandbox network policy blocked access to the World Bank data file. Claude stopped and asked whether I wanted to upload the page (as a pdf) myself or let the work continue without it. I skipped the question, so it continued and took its information from the World Bank’s GPPD DataBank API instead. However, the DataBank holds 2018 data, while the portal shows a 2022 profile. As a result, Claude’s baseline data was different from GPT and Muse.
Muse did not ask for any permission until the fourth part of the task, which required registering on the World Bank website and uploading information.
Account registration on the World Bank website
The final part of the task, registering on the World Bank portal and uploading the changes, is where things became more interesting because it required the agents to take a more significant actions rather than just gather information.
Claude and GPT both stopped at this point and handed the registration and uploading over to me. Muse, however, kept going. Without asking me or showing me the Terms of Service, it registered an account under the email address [email protected]. You can see Muse’s full back and forth here.
The table below summarizes how each agent approached this last part of the task and cybersceuirty implications about it.1
Monitoring and observability agents vary widely
There is an ongoing debate about whether AI labs should expose a model’s full chain of thought (CoT) and action trace, and if so, how much. Labs have given several reasons for holding back. OpenAI chose not to show o1’s raw CoT to users, citing user experience, competitive advantage, and the value of keeping the CoT available for internal monitoring. Anthropic noted that raw reasoning can contain incorrect or half-formed thoughts and that malicious actors could use it to build better jailbreaks. There is also a gaming and reward hacking concern, and “CoT unfaithfulness”.
To understand an agent’s behavior, however, evaluators need to know when and why things happen, which is only possible with a monitoring system in place and access to the agent’s complete trajectory. For an evaluator outside an AI lab, without that access, it is nearly impossible to fully make sense of an agent’s behavior. And if outside evaluators can only see partial trajectories, and any conclusions they draw can be dismissed for lacking complete information, what is the value of independent evaluation?
Knowing these limitations, I tried my best to collect, monitor, and check as much of each agent’s work as I could, again putting myself in the position of an ordinary researcher tasked with updating the World Bank information portal.
Multilingual performance the agents could write in Farsi better
A few observations and then I’ll get to my points:
- All three agents answered in fluent Farsi, but the Iran results were far weaker than the US ones. Of Iran’s 138 N/A fields, GPT and Muse each filled only 21 with a real value; for the US, they filled 51 and 64 of 130.
- For the US, 76–89% of each agent’s citations came from official government sites and the rest from legitimate international organization websites. For Iran, it was 11–22%.
- Low-authority sources crept in, including a Telegram channel, a Medium post, Grokipedia, or websites run by Iranian diaspora media groups such as Iran International. Claude seemed to be more conservative about finding workarounds when websites were unavailable and often preferred English language sources even with low legitimacy.
- Claude could only open 3 out of the 16 Farsi pages it tried. GPT and Muse cited 11 Persian sources each but showed reading only 3 and 7 of them respectively.
Knowledge gaps got filled with something else: Claude used headlines and its own memory (sometimes contradicting with what it found), and said so. GPT used republished copies of the law. Muse mostly read a 2009 English translation but cited the official Persian page.
My point is not that I expected the Iran/Farsi tasks to have the same outcomes as the US/English ones. After all, the Iranian government has made it very difficult for foreign IP addresses to access official websites and domains ending in .ir (you can read more about this in the context of Iran’s National Information Network). My point is about the agents’ differing workarounds and source prioritization.
For me, this brought to mind the digital rights and language inclusion work that the good people of Global Voices have done for years, including on net neutrality and language access. What does all this mean for an AI agents era? And from an AI sovereignty perspective? One of AI sovereignty’s promises has been language diversity and support for local languages. LLM output quality, and perhaps safeguards, keep improving, but we also need to think about what language localization should look like in agents reasoning, searching, and prioritizing sources.
If you are interested in designing or conducting experiments on this topic, feel free to reach out at [email protected].
Share the article:
Trending News & Crypto World
Latest Crypto News
Top Trending Cryptocurrencies on The Market
Current Price
$0.9999
Market Cap
$2.9B -0.05%
24h Volume
$161.7M
Supplies
2.9B / ∞
Current Price
$178.36
Market Cap
$2.8B 0.10%
24h Volume
$579.7M
Supplies
16.0M / 16.0M
Current Price
$0.005490
Market Cap
$2.6B -6.29%
24h Volume
$603.0M
Supplies
830.3B / 1.0T
Current Price
$120.30
Market Cap
$2.5B -1.17%
24h Volume
$20.7M
Supplies
21.0M / 21.0M
Current Price
$1.000
Market Cap
$2.5B -0.25%
24h Volume
$254.7M
Supplies
2.5B / ∞
Current Price
$1.140
Market Cap
$2.4B 0.03%
24h Volume
$-
Supplies
2.1B / ∞
Current Price
$1.050
Market Cap
$2.4B -0.34%
24h Volume
$1.1M
Supplies
5.4B / 10.0B
Current Price
$0.4828
Market Cap
$2.4B -4.14%
24h Volume
$246.5M
Supplies
10.0B / 10.0B
Current Price
$0.2313
Market Cap
$2.3B -5.40%
24h Volume
$460.8M
Supplies
15.0B / 15.0B
Current Price
$1.150
Market Cap
$2.3B 0.29%
24h Volume
$1.3M
Supplies
2.0B / ∞
Current Price
$1.000
Market Cap
$2.3B -0.02%
24h Volume
$-
Supplies
2.3B / ∞
Current Price
$0.6509
Market Cap
$2.1B -4.46%
24h Volume
$31.9M
Supplies
6.2B / 6.2B