Avoiding AI Bias
We’ve been building really efficient IT systems for a long time. But that efficiency comes at a cost – often the cost of context. It turns out that we’ve spend decades building reams of transactional data, efficiently stored without the heft of all the contextual data that would give it meaning.
And now context is the new hot thing. If you want meaning, you need contextual data, stored somewhere that’s machine readable, i.e., not in your head (at least not yet). In good news, tools like MCP (Model Context Protocol) and semantic layers help overcome data silos – but they can’t mitigate for data that’s not captured in the first place.
It’s true that organizations have oceans of data. But lots of that data is irrelevant, or uninterpretable. And this isn’t even a new problem – lots of folks have been making noise about this for over a decade, and we’ve collectively just been kicking the context can down the road for all that time.
The things generative AI and Agents can do are nothing short of amazing. But they are limited by the data they’ve got. I’m really excited to see organizations layering in their own data sets for more relevant results from these kinds of engines, but I really worry about the quality of the datasets being used.
One of the famous examples is the tendency of AI’s to label any kind of windswept, treeless landscape as “sheep.” Those engines successfully pattern matched the image as a whole, but they attached the label to the wrong features.
If there are distinctive features, it’s reasonable to expect that those distinctive features will be the ones that AI’s identify. But what happens when the distinctive features are irrelevant to the outcome variable of interest… and introduce bias into the decision-making process?
One consequence is that, by (inappropriately) excluding some features (or people) from considerations, you’re necessarily reducing the overall quality of your overall results.
If you’re trying to find the top 10% of all Sneeches (of Dr. Suess fame), but you only consider Star Bellied Sneeches, you’ll likely end up selecting a lot of Star Bellied Sneeches that perform worse than the best Plain Sneeches would have performed. You’ve lowered your average performance, and for no good reason.
I worry a lot about this as we hand over more decision making to AI’s and agents. Automation strategies like RPA (Robotic process optimization) are terrific for efficiency and their outputs are constrained. Narrow application, but low risk.
But the seductive power of Gen AI and Agents are that they can handle novel inputs and generate novel outputs, and handle enormous numbers of transactions autonomously. There’s been more AI than human generated content on the internet for over a year now, leading to risks of AI model collapse.
One of the great promises of handing over decisions to machines was that they could make better decisions – more pure, more unbiased. And, that’s definitely possible. But it generally requires very thoughtful curation of training datasets to be as uncontaminated by biases of any flavor as possible. Current models are so data hungry, it seems as if many are just shoveling in any data we can get our hands on, precision be damned.
And, as AI-generated content is used to feed AI models, all the little warts and wobbles get amplified. That’s referred to as recursive training and it’s the data equivalent of inbreeding. Enough generations of this and you end up ruled by feeble monarchs who are so deformed they are barely able to chew.
We’re not doomed; we know how to do better. But doing better requires intention and effort. Here’s how to build your AI ecosystem to help build the world we want, rather than reproducing the warty world we’ve got:
1. Audit your training datasets. Training data represents the version of the world that the system will try to match. Think of this as presenting the best version of yourself on a first date. Put your best self out there!
2. Invest in robust governance. Like semantic layers. Guide your systems by prescribing meaning and limiting irrelevant data.
3. Build output guardrails. Build continuous monitoring to flag problems before they get out of control and audit regularly.
AI systems are amplifiers – and they’ll give us back a refined version of what we fed them. If we are thoughtful, intentional, and attentive, we can build systems that help us live into the best versions of ourselves. But if we are seduced by shortcuts and haphazard approaches, we’ll find our worst impulses amplified – faster results, but worse outcomes.

This is such an important reminder: AI doesn’t just amplify capability, it amplifies assumptions. Without context and careful curation, efficiency turns into bias at scale.