Perch Insights
Perch vs. a general AI, both on your data

What actually delivers a spot-on insight?

AI analytics is finally real, and everyone now claims you can point an LLM at your data and get answers. So we ran the comparison that matters. We put Perch next to a capable general AI, both pointed at the same real program data, and looked at one thing: which one produces an answer your team sits with, checks against what they already know, and calls spot on.

That reaction, not a benchmark score, is what decides whether an insight gets acted on or quietly ignored. A number can be accurate and still get waved off in the room. The answer that gets used is the one a leader trusts enough to act on.

“Yes. That makes sense.”
“Ugh. That misses the point.”
01 · The gut check

What “spot on” actually means.

You cannot always name why an answer is good. You just feel it. When we look closely, that feeling comes down to a handful of moments, and every one of them has a version that lands and a version that makes you close the tab. This is the vocabulary of trust, in the words your team actually uses.

It understands my businessDoes it know the motion behind the metric?
Spot on
It knows our funnel runs lead to contact to opportunity to sale, and that contact rate is where the outbound motion lives or dies.
×Misses the mark
It reads the program like a generic sales report. It never connects to the outbound motion, so it does not even see contact rate as part of the story.
It gets to the real root causeDoes it triangulate, or stop at what moved?
Spot on
It goes past what moved. It checks agent execution against call-quality scores, weighs that against lead quality, and keeps going until it can say what is actually responsible.
×Misses the mark
It names what moved and stops. Then it hands me a list of things to go investigate, which was the whole part I needed it to do.
It compares like with likeSame-length period, fair benchmark?
Spot on
Like for like on time. It does not stack our short February against January as if they are the same, and it reads the trend on the year-to-date view we actually run on.
×Misses the mark
It compared February straight to January and called the dip a trend, ignoring that the months are not even the same length.
It picks the right metric, cut the right wayRight definition, fair cohort?
Spot on
It used the lead-journey cohort contact rate for this question, not the raw per-call one, and it compared tenured agents only against tenured agents. The right measure, filtered to a fair comparison.
×Misses the mark
It grabbed whatever “contact rate” it found first, and lumped week-old hires in with veterans, so the comparison never meant anything.
It stays consistentSame method, every time you ask?
Spot on
Ask the same question today or next month, and it runs it the same way, same definitions, same exclusions. The number holds up in the meeting.
×Misses the mark
It normalized per agent one week and forgot the next. Same question, different method, and now nobody trusts the number.
Its recommendations are actually doableA move that survives how we operate?
Spot on
It gives moves that survive how we actually run. It would not tell me to cut a lead source without saying what backfills the volume, or the agents go idle and we miss the number.
×Misses the mark
It told us to stop dialing after attempt 7 to save effort. But there are real conversions past 7 and the agents have the capacity. Following that just hands away sales.

Watch for these six in the two questions that follow. They are the whole difference between a nod and a shrug.

02 · Two real questions

The same questions, read by the person who has to act on them.

We compared two things: Perch on your data, and a capable general AI on your data. Same program, same prompts. For each question we show the strong, spot-on answer first, then read both the way your team would.

Perch
Running on its complete, connected model
A capable general AI
Pointed at the same data
1The open questionNo guidance given
Prompt to both systems“Walk me through what happened in the program last month.”
Perch

Total sales fell 18.4%, driven mainly by a contact rate collapse, 56.0% down to 48.2%, with the rest downstream. Agent effort was flat: calls per agent per day did not move.

The drop is uniform across all 10 states, all 17 vendors, every hour of the day, and even the earliest attempts. That signature points to spam-labeled outbound numbers, not lead quality.

Recommended action · two movesFirst, specific outbound numbers have fallen off a cliff and need to be rotated now. Second, last period we stopped dialing many leads short of the full 10 attempts we commit to, so reachable leads are being left on the table.
A capable general AI

Sold-lead volume fell 22.7% month over month. The top of the funnel looks like the bottleneck: opportunities worked dropped 10.7%, and interactions per sold lead rose 21%, so agents are working harder per sale.

This reads as a lead-supply or lead-quality softening. Ticket size and bundle rate held.

Recommended actionWork with marketing and lead ops on why opportunities-worked fell, and review the lower-performing producers.
Reading Perch
“Yes, that makes sense. Our own outbound motion is the problem, and it told me which numbers to rotate and where we stopped dialing early.”
×Reading the AI
“Ugh. It did not even lead with what happened to sales, and it never clocked that our outbound motion, contact rate, is a huge part of this funnel. It is sending me to marketing without checking whether our own contact strategy is the real problem.”
2The guided questionCuts handed to the AI
Prompt to both systems“Close rate dropped in February. Dig in, and look at agents, lead sources, and states.”
Perch

Close rate fell 100 bps to 16.1%. Your three cuts, plus what they open up:

Agents, done like for like: before comparing producers, I separated the 8 agents who started in January from the tenured team. Much of the apparent agent decline is just new hires ramping. On a tenured-only basis the real drop is smaller, and points elsewhere. Lead sources, with the link you did not ask for: the vendor decline likely traces to last month’s spam-labeling finding. If flagging hit those vendors hardest, the contacts getting through are less interested, which reads as a lower close rate. The lever nobody flagged: 30% of open opportunities got no follow-up attempt at all. Bigger and more fixable than the vendor mix.

Recommended actionEnforce a follow-up SLA first, hold new hires to a ramp curve instead of the veteran benchmark, then coach the cost and timing objection dip.
A capable general AI

Close rate fell from 17.6% to 15.4%. Across the three cuts you asked for:

Agents: the decline concentrates in a few producers; a couple actually improved. Lead sources: about two-thirds of the drop traces to four vendors, with an estimated sales gap for each. States: NE and OH fell hardest.

Recommended actionStart with the four underperforming vendors. The agent and state effects look secondary.
Notice the two systems do not even agree on what close rate did. Perch runs one governed definition. The general AI reached for whichever version it found first, which is exactly why the number will move when we ask the same question again. More on that below.
Reading Perch
“Yes, that makes sense, and it is what I would have asked for if I had thought of it. It caught that my new hires were dragging the whole average, and it found a follow-up hole I never mentioned.”
Reading the AI
“Not bad. It answered the three cuts I gave it, and the vendor number is concrete.”
Question 2, one week later
We asked both systems the exact same question again. Watch what changed.
Perch
The same answer. Same lead-journey cohort contact rate, same tenured-versus-new split, same follow-up branch. The method did not move, so the number did not move.
A capable general AI
A different answer. This time it used the per-call contact rate, dropped the tenure split, and pinned most of the drop on states. Same question, one week later, a new story.
“So which number do I bring to the Monday meeting?”
Hand a capable AI the exact cuts to run, and it does well, once. Ask it an open question about your business, or the same question twice, and it drifts. Perch was spot on every time, because the analysis you would otherwise have to spell out is already built in.
Notice what carried Perch on question two: it went past the cuts you asked for, and it knew to compare new agents against new agents. No larger model hands you that. Knowing your business before you open the tool does.
A note on the setup: on question two we deliberately handed the general AI the specific cuts to run (agents, lead sources, states), the fair version of the test, and it did well within them. On the open question, and on the repeat a week later, it was on its own. Perch received no such help on any of the three.
03 · Why the reactions are so different

Same model. The difference is what it can see.

An AI is only ever as good as the picture it reasons over. Perch reasons over a clean one: every call joined to the lead it was made against, one definition of every metric, the business itself encoded, who is new, what the cap is, what to leave out.

Point the same capable AI at your data as it sits today, and it sees noise: calls that do not connect to leads, three definitions of close rate, no idea which agents are new, test records mixed in with real ones. Same reasoning. Opposite result.

✓  The AI on the Perch model
every call joined to its lead one definition, everywhere tenure and the 10-attempt protocol known noise removed
What comes out
“Sales fell because contact rate collapsed from 56.0% to 48.2%, uniformly across every state and vendor, the signature of spam-labeled outbound numbers. Effort was flat. Rotate the flagged numbers and finish the attempt cycle.”
“Yes, that makes sense. And I know what to do.
Take that model away, point the same AI at your data as it sits, and the only thing that changed is what it can see.
×  The same AI on your raw data
calls not linked to leads three definitions of close rate new hires mixed with veterans test and recycled leads included
What comes out
“Sold-lead volume is down and interactions per sale are up. Recommend reviewing lead quality and coaching the bottom-quartile agents.”
“Ugh. Plausible, confident, and pointed at the wrong problem.

So the answers above are not different because one model is smarter. They are different because of the layers Perch puts underneath the reasoning. Here is what each layer does, and the exact difference it made on the very answers you just read.

A connected data model
Joins your systems so every call ties to the exact lead it was made against, and clears out the test and junk records.
Why the answers differedWithout it, the general AI reasoned over unrelated rows and read noise as signal.
Governed definitions
One agreed meaning for every metric, and the right version chosen for each question.
Why the answers differedThis is why Perch held to the lead-journey cohort contact rate, while the AI switched to a per-call definition a week later.
The diagnostic playbook
The map of what drives each metric, and the order to check the drivers, built from years of running these programs.
Why the answers differedThis is why Perch triangulated agents against lead quality and went past the cuts you asked for.
Your business, encoded
Your funnel, your 10-attempt protocol, who is tenured, and what to leave out.
Why the answers differedThis is why Perch read a reachability problem as reachability, and compared like with like.
Cross-program learning
Patterns found across many programs feed back into the playbook.
Why the answers differedThis is why the diagnostics keep sharpening instead of starting from scratch on every question.
The actual product

Perch builds these layers and, the harder part, keeps them current as campaigns launch, dialers migrate, agents turn over, and product mix shifts. That is what you are buying: not a cleverer model, but a living system of knowledge and structure, maintained so that the AI reasoning on top of it produces spot-on insights continuously, not just on a good day.

21%95%+
Independent research reached the same conclusion. A capable model pointed at a data corpus hit 21% accuracy. The same model, given a curated, structured knowledge layer, cleared 95%, and near 99% in some domains. Handing it more raw data access moved the number less than a single point. The bottleneck was structure, the exact thing a general setup on raw data does not have.Source: Anthropic data science team, “How Anthropic enables self-service data analytics with Claude,” June 2026.
In one line
A capable AI can be brilliant on a good day, on a question you framed, over data you cleaned. Perch is spot on by default.

It answers the questions you did not know to ask, over a model that already understands your business, the same way every time you open it. Your team will not grade it on a benchmark. They will grade it on whether the answer makes sense, and that reaction is the whole product.

“Yes, that makes sense.”
“Ugh, close but no.”
The comparison that counts

The one that matters runs on your data.

The head-to-head on this page ran on real program data. Give us your last ninety days and one question you have never gotten a clean answer to, and we will show you what spot on looks like on your own numbers.