In a recent LinkedIn post, Adam Kucharski perfectly illustrated the big risk companies are facing when they adopt AI tools to “scale productivity” in the data analysis domain:
I gave Claude Code a real-life behavioural dataset and asked whether there were any interesting patterns. A few minutes and a few thousand tokens later, it had a clear answer for me:
“I ran an exploratory analysis across all pairwise correlations and group comparisons in the 100-person behavioural dataset. The strongest statistically significant finding was: Higher education level is associated with fewer monthly leisure activities (Spearman rho = -0.23, p = 0.019). […]”
As Kucharski says, at first glance, the fully automated data analysis and confident conclusion look impressive.
Yet, his dataset had been created by randomly simulating 100 individuals with 5 demographic characteristics and 6 behavioral indicators, which means that none of the AI conclusions had merit. As a statistically literate person, Kucharski wasn’t fooled, but the technically polished but substantively bogus AI output could have easily mislead an unsuspecting bystander.
A qualified analyst would have taken much longer than AI to complete the analysis, but they would have never made this statistical mistake when performing the task.
The fact that the dataset was based on “random noise” doesn’t even matter here (except as a way to efficiently prove the point about the AI analysis being wrong). Even when data collection is carefully designed and executed, issues like p-hacking can still occur during analysis and reporting. The problem here isn’t how the data was created, but what happens after the data exists.
AI enthusiasts will say that this kind of mistake can be prevented by giving AI “more context” and guardrails, as well as by reviewing the output for technical flaws and asking for correction.
Well, what they’re saying then is that generic AI tools can only be reliably used by experienced professionals leveraging it to accelerate their results without delegating the actual thinking.
But at the moment we have lots of CEOs (I know some of them around the world, including in my native country, Brazil) celebrating how they’re being able to reduce their workforce using AI to “scale productivity” by replacing seasoned (read: expensive) programmers and data analysts with a team of much cheaper junior professionals armed with generic AI tools.
I thank them when they boast publicly about adopting AI while going through rounds of layoffs involving senior staff. That means I can mitigate my investment exposure to businesses facing the elevated risks of delegating tasks like statistical analysis or programming to a generic AI tool without robust supervision. Sooner or later they’ll end up developing false confidence in incorrect data findings and/or having to deal with potentially disastrous bugs and security vulnerabilities in their internal software.
The risks can be significantly reduced by choosing dedicated AI tools that have not only been built with robust guardrails to perform a specific task, but fully tested and vetted for the specific context of the business, data, and codebase.
And on top of that, one has to make sure that there are enough experienced professionals around to proactively anticipate and detect errors before they cause any harm.
But who wants to go through such heavy investments in technology and human expertise when hyperbolic statements by tech leaders insist that giving novice users access to generic AI tools will aggregate into huge “productivity gains”? Perhaps it will happen once the damage of a flawed AI output that is public-facing or tied to sensitive decisions starts to cascade across reputation, legal exposure, and day-to-day operations.


