The person saying “AI is incredible” is not hallucinating. The person saying “AI is unreliable” is not being a hater. They are likely just touching different parts of the frontier. — Corey Noles
A lot has been written recently about the concept of “jagged AI frontier”. If you aren’t familiar with this expression, MIT Sloan’s working definition and Tomas Pueyo’s viral image below can help get you up-to-speed:

We’ve all seen many examples of how AI, despite dramatically surpassing human cognition in many areas, can still falter in simple tasks. But there’s another layer to the problem of the jagged frontier that makes it even trickier to manage: its non-monotonicity.
The jaggedness of AI progress doesn’t move in a single direction. It’s not simply a matter of weak areas catching up over time — sometimes capabilities that were previously strong can deteriorate, even as others advance.
No one would expect a player who used to excel at both checkers and chess to continue to improve their chess game while starting to lose in games of checkers, right? But with AI models, it’s not uncommon for a replacement model (from the same AI lab, with better benchmark scores than the previous version), to start giving incorrect answers, or making inference mistakes, in contexts where it used to work well before.

People with deep knowledge in a domain where they use AI see examples of the “jagged AI frontier”all the time, even in interactions with the most powerful models.
AI labs regularly update their models, and this is when the dynamic nature of the jagged AI frontier may cause the most unexpected failures.
For example, a colleague recently told me about an issue he experienced with a paid model after its latest update. He asked the AI whether a higher RMSE is good or bad when evaluating a predictive model. (RMSE, or root mean square error, is a common measure of how far a model’s predicted values are from the actual values. A lower RMSE means the model fits better.) To his surprise, the model “confidently” said that a higher RMSE is better—a mistake the previous version of the same AI model would not have made.
This is why companies adopting AI agents to automate their workflows need to worry not just about the unexpected weaknesses their AI models may have today, but also about new potential weaknesses they may develop in the future.
For instance, imagine a healthcare company implementing an AI-based self-service solution that interacts directly with healthcare providers to help them get their medical claims approved.
The company does its due diligence, defining what good behavior looks like and turning domain expertise from its claims analysts into concrete, testable criteria that engineers and operations teams can use to determine whether the solution is ready to ship. As time passes, the AI service is given more autonomy based on its record of reliability.
At this point, if the work was still being done by experienced claims analyst, the company wouldn’t have to worry about them suddenly starting to make basic mistakes, such as divulging patient confidential data or making wrong calculations.
But with AI models, that risk is there, especially when a model update may suddenly change the existing jagged frontiers in a way that break parts of a tested workflow.
In this scenario, incorrect approvals could facilitate improper payments or fraud; a data breach involving medical claims data could expose the company to lawsuits or other causes of action; failure to meet claims-processing service levels could violate agreements, etc. And with thousands and thousands of claims being processed at a much higher speed than analysts can oversee, it might take time to understand what’s happening across an entire body of transactions.
So far I’ve been talking about the downsides of the jagged edge frontier, but in the title of this article I mention that I also see a benefit to it.
And the positive side I can see in the jagged AI frontier is that it helps explain why companies doing AI-driven layoffs are often proven wrong, with costly reversals already happening.
AI may be dramatically changing how work gets done, but we still need domain experts with deep knowledge of the workflows, policies, and edge cases the AI agents are tasked with handing; professionals who can turn this knowledge into prompts with precise instructions to be consistently followed; teams capable of developing sophisticated evaluation processes that run daily; and so forth.
Of course, this “positive side” that I’m celebrating will only benefit the professionals who have or can develop the skills that are becoming more and more valuable across industries.
Skills like technical depth, analytical fluency, human judgment, coordination, domain expertise, are going to remain in demand. But what does it mean for early-career professionals who didn’t yet have a chance to develop such capabilities? Time will tell if employers will find ways to help juniors become seniors in an environment where the opportunities to develop our intellect by working on problems that now AI can easily solve for us have become increasingly rare.


