GenAI Testing: How to Find Idiom and Slang Failures in AI Models
Introduction
On a project in Jakarta, one of our crowd members typed a prompt into an AI tool in Bahasa Indonesia, using the phrasing and slang she'd use day to day. The model got every word right, until she used an idiom. Gulung tikar literally means rolling up the mat, and it's how you say a business has gone under. The tool took her at face value and answered as though she'd asked about an actual mat.
It wasn't a translation error. Every word was correct. The tool simply had no idea what she meant.
This is episode two of Field Notes, a series where we take a closer look at the small details that decide whether software actually works for the people using it. Today: what happens when GenAI learns a language without learning how people use it.
Field Notes, Episode 2: GenAI — 1min 24 sec. Full transcript at the bottom of this page.
Fluent Is Not the Same as Fluent in Their Words
We've been seeing this across a lot of our GenAI customers. These tools become genuinely fluent in a local language without ever picking up on how people actually use it.
What gets lost is the layer sitting on top of vocabulary and grammar: idioms, like the one above; slang, which shifts by city and by generation; and register, the signals that tell you how formal someone is being or what kind of mood they're in. A user who switches to clipped, informal phrasing is telling you something. A user being elaborately polite is telling you something else. A model that reads both as neutral prose is missing most of the conversation.
The failure mode is specific and easy to underestimate. The tool responds confidently, correctly at the word level, and completely beside the point.
Why This Is Almost Invisible From the Inside
It's extremely difficult to spot that your tool is reading people this way, because a misunderstanding like that won't necessarily surface in your performance data.
Nothing errored. Latency was fine. The response was well-formed, grammatical, and on-topic by any automated measure. Your evaluation metrics have no way to register that the answer addressed a floor covering rather than a bankruptcy.
The only person who knows the tool got it wrong is the one who asked. And she won't file a ticket. She'll decide the tool doesn't really get her, and stop coming back.
That's the pattern that makes this expensive. You don't get a bug report. You get a slow decline in engagement in one market, which is very easy to explain away as lower product-market fit rather than a quality problem you could actually fix.
Why GenAI Testing Needs People, Not Just Benchmarks
Standard evaluation catches the things you can score. Did the model produce grammatical output? Did it stay on topic? Did it avoid an obviously unsafe answer? Those are all worth measuring, and none of them detect an idiom taken literally.
Catching that requires someone who speaks the language the way your users speak it, not a textbook version of it, and who is free to probe rather than follow a script. It's closer to exploratory testing than to a benchmark run: a native speaker deliberately reaching for the slang, the idioms, the sarcasm and the shifts in formality that a real user would produce without thinking about it, then judging whether the model actually understood.
That's the work we do. Global App Testing puts software in front of testers who live in the markets our customers are targeting, using the language as it's actually spoken there, so the misreadings that only a local speaker would notice stop being invisible to teams who aren't there. You can see how that works on our people-powered testing platform.
The Question to Take Back to Your Team
Your GenAI can hold a conversation in your users' language. But has anyone checked whether it can hold one in their words?
Worth asking specifically, per market, and worth asking someone who would actually notice the difference.
Full Transcript
Veronica, Marketing Manager at Global App Testing:
Hi, my name is Veronica, and this is Field Notes, where we take a closer look at the details that shape whether software actually works for the people using it.
In this episode, I wanted to share with you something we've been seeing with a lot of our GenAI customers.
Often, these tools are becoming fluent in a local language without ever picking up on how people actually use it: the idioms and the slang that tell you how formal someone is, or what kind of mood they're in.
For example, on a project in Jakarta, one of our crowd members typed a prompt into an AI tool in Bahasa Indonesia, using the phrasing and slang she'd use day to day. The model got every word right. But the second she used an idiom, gulung tikar, which literally means rolling up the mat and is how you'd say a business has gone under, the tool took her at face value and answered as though she'd asked about an actual mat.
And it's extremely difficult to spot that your tool is reading people this way, because a misunderstanding like that won't necessarily surface in noisy performance data. The only person who knows the tool is getting it wrong is the one who asked, and she won't file a ticket. She'll just decide the tool doesn't really get her and stop coming back.
So here's the question I'd take back to your team. Your GenAI can hold a conversation in your users' language, but has anyone checked whether it can hold one in their words?
Conclusion
A model that misreads an idiom isn't broken in any way your dashboard can see. It passes your evals, returns clean output, and quietly loses the user who noticed.
Finding those moments means testing with people who speak the language your users actually speak. Talk to us about putting your GenAI in front of native speakers in your target markets, and find out what your model sounds like to the people you built it for.
Field Notes is a series from Global App Testing on the details that determine whether software works in the real world. Catch up on Episode 1, on payment methods.
Validate Real-World Language Understanding Before Your Users Do
Ensure your GenAI handles the idioms, slang, and shifts in tone your users actually use, with independent testing from native speakers in your target markets.