The Problem-Selection Problem: When Agents Choose What to Automate
What happens when you give two AI agents open access to the same company's files and Slack and ask them to find the problem worth solving? They don't agree, and that gap is where most agentic projects actually go wrong.

Give an AI agent full access to your files and Slack, tell it nothing about what to automate, and ask it to come back with a problem worth solving. Two different agents did exactly that against the same business recently, and they picked two different problems entirely. That gap is the real story.
Why would two AI systems disagree about what's broken?
The setup was simple: open access, no assigned task, one instruction, find a real problem and build something for it. One agent came back with a narrow, already-articulated fix: make the internal handoff package cleaner so the next stage of work could start faster. It was a real problem, it was finished in a single run with no issues, and it was, by the account of the person who ran the test, "fine", something they'd probably use.
The other agent went somewhere nobody had asked it to go. Instead of the handoff itself, it identified the harder problem sitting one level upstream: deciding which idea was worth pursuing in the first place. Nobody had put that problem into words before the agent named it. The resulting tool wasn't rated "fine", it was called essential, something the person couldn't now do without.
Same files. Same Slack history. Same company. Two agents, two different diagnoses. One stayed inside the boundary of a problem that had already been spoken out loud somewhere in the business. The other went looking for the problem underneath the problem.
The gap between what people say and what they do
This tracks with something we keep noticing: most people have a different verbal account of their business's problems than their actual behavior shows. Ask someone where the friction is and you'll get a tidy answer, usually whatever was most recently annoying, or whatever's easiest to describe in a sentence. Watch what they actually do, what piles up unread, what gets rerouted three times before it's handled, and a different picture emerges.
It's the same reason 1.6 million agents were registered for one widely hyped agent platform and most never completed a single task. Capability wasn't the bottleneck. Knowing what to point it at was.
Why do so many agentic projects stall anyway?
Gartner estimates more than 40% of agentic AI projects will be shut down by the end of 2027, citing unclear business value and inadequate controls as leading causes, not immature technology. A pattern behind a lot of those failures: teams describe their problem once, broadly, and every vendor or agent hears something different in it. An accounts receivable team doesn't have one AI problem, collections prioritization, invoice matching, customer follow-up, exception handling, cash application, dispute resolution, reporting, and escalation are each a different shape of work, and a solution built for one rarely fits the rest. Fold them all into a single ask and you get a tool that's mediocre everywhere.
What this means before you automate anything
The instinct is to jump straight to picking a tool. The harder, more useful step comes first: let something look at how the work actually happens (the files, the threads, the half-finished folders) rather than relying only on how people describe it in a meeting. Stakeholder interviews aren't wrong, exactly; they're just partial. They capture what's top of mind, not what's actually costing the most time.
| Starting point | What it tends to surface |
|---|---|
| Ask the team what's broken | The most recent or most describable annoyance |
| Read the actual work artifacts | The pattern nobody had bothered to name |
Neither approach alone is complete. But the second one is the piece most businesses skip entirely, and it's exactly why identical raw material can produce two entirely different, and differently valuable, answers.
FAQ
If two agents can pick different problems from the same data, how do you know which one is right? You probably don't, on the first try. Treat the first pass as a hypothesis, not a verdict, check it against what a person closest to the work actually recognizes as painful, and be willing to run it more than once.
Doesn't this just mean agents are unreliable? Not quite, it means problem-finding is a judgment call, not a lookup. The same way two consultants can walk the same floor and flag different risks, agents reasoning over open-ended material will emphasize different things. That's a reason to treat problem selection as its own step, not a reason to skip agents altogether.
Before you invest in building or buying automation, it might be worth spending a little time somewhere else first: figuring out, with real evidence rather than a hallway conversation, what's actually worth automating.
More on AI Agents
Want a system like this in your business?
We build the automation behind everything you just read.


