Stop Creating AI Slop: How to Build AI Agents That Create Value for Customers - Botco.ai

Stop Creating AI Slop: How to Build AI Agents That Create Value for Customers

AI can sound remarkably capable.

It can answer a question fluently. Summarize information. Hold a natural conversation. Produce a polished response in seconds.

But none of those things necessarily mean it created value.

That was the central discussion in Botco.ai’s recent webinar, “Stop Creating AI Slop: How to Build AI Agents That Create Value for Customers,” featuring Rebecca Clyde, CEO of Botco.ai, and Chris Maeda, CTO of Botco.ai.

Their message was simple: a fluent AI agent is not necessarily a useful AI agent.

An enterprise agent has to do more than generate a convincing response. It needs the right information, access to the right systems, clearly defined permissions and workflows, guardrails around what it can and cannot do, and a way to learn from the interactions that don’t go as planned.

The webinar broke down what that looks like in practice.

1. AI slop can sound perfectly intelligent

AI slop isn’t always obviously wrong.

Sometimes it’s a long answer to a simple question. Sometimes it repeats information the audience already understands. Sometimes it gives generic guidance rather than a specific answer.

And in an enterprise setting, it can be even more expensive: the AI answers the customer’s question but leaves the customer to finish the task themselves.

As Rebecca explained:

“One of the things that often will characterize AI slop is that it sounds very confident.” — Rebecca Clyde

Chris described encountering the same problem while using AI to produce technical release notes. The first result was verbose and explained concepts that the intended technical audience already understood. After receiving specific feedback about the audience and desired response, the model rewrote the material using substantially fewer words.

That illustrates an important distinction:

More output doesn’t mean more value.

The goal is not to make AI say more. It is to make AI provide the right information, at the right level of detail, for the person and task in front of it.

2. Answering is not the same as resolving

Consider a customer asking:

“Can you reschedule my appointment?”

An AI system could respond at three very different levels:

Answer: Tell the customer how to reschedule.

Assist: Find available appointment times.

Resolve: Find availability, reschedule the appointment in the appropriate system and confirm the change.

Only the third completes the customer’s intended outcome. The webinar deck captured this distinction succinctly: “Measure outcomes, not conversations.”

Rebecca explained why this changes the way organizations should think about AI performance:

“We’re not just measuring a conversation took place… we want to be able to complete that whole procedure.” — Rebecca Clyde

Resolution requires more than a language model.

Behind a seemingly simple conversation may sit identity and permissions, workflow logic, guardrails, an EHR, CRM or scheduling platform, and an action performed against one or more of those systems. The conversation is simply the front end; the architecture behind it determines whether the agent can actually finish the job.

3. Start with the outcome—not the AI

One of the most useful reframes from the session was also one of the simplest:

Don’t start with: “Where can we add AI?”

Start with: “What customer outcome are we trying to complete?”

Once the outcome is defined, organizations can work backward.

What information does the agent need? Which systems contain it? What actions must the agent perform? What permissions are required? Where should a human become involved?

This is where integration stops being merely an IT discussion and becomes part of the customer experience.

An agent that cannot reach the systems where the work happens cannot reliably complete the work.

For healthcare organizations, for example, that might mean connecting an agent to an EHR for relevant patient information, a CRM for customer context, a scheduling system for availability, a CMS for validated content or a contact center for escalation and human handoff.

4. The leap from demo to production is bigger than it looks

A polished demo proves that an AI agent can successfully complete a scenario.

Production requires something much harder.

It has to deal with the scenarios you didn’t rehearse.

When an audience member asked what typically takes the most work moving an AI agent from demo to production, Chris summarized the difference:

“A demo is… let me show you one thing working correctly.” — Chris Maeda

Production, he explained, means supporting the breadth of situations customers may encounter—or being able to fail gracefully when the system cannot complete something.

That includes incomplete information, unexpected answers, multiple intents, unsupported requests, system failures and escalation paths.

This is why production testing cannot stop at the happy path.

The Botco.ai Playground is designed around this idea: teams can refine agents before customers encounter them. The principle highlighted during the webinar was particularly useful:

Don’t only test how it answers. Test how it fails.

5. Training an AI agent is really a feedback loop

One of the most practical parts of the discussion concerned what “training” an enterprise AI agent actually means today.

Organizations sometimes assume they need to retrain or fine-tune a large language model every time an agent produces a poor response.

Chris explained that, in many business applications, improvement increasingly happens through better data, better prompts and better examples.

“The more concrete you can make the feedback to the models, the better they perform on any given task.” — Chris Maeda

That feedback can include positive examples, negative examples, corrected answers and human review.

For example, if an AI response is unnecessarily verbose, tell it what information the intended audience already understands. If it handles a particular question incorrectly, provide an example of the appropriate response. If a particular request should never be handled by AI, explicitly define that boundary.

Botco.ai supports these feedback loops by allowing teams to review responses and revise AI answers, helping agents improve how they handle similar interactions. The webinar deck makes feedback loops one of the eight core checks for an enterprise-ready agent.

6. In healthcare, “Can the AI do it?” is the wrong question

Healthcare raises the stakes considerably.

The more useful question is:

Should the AI do it?

An agent may technically be capable of generating an answer, accessing information or initiating an action. That doesn’t mean it should have permission to do so.

For every workflow, teams need to determine:

What can the agent see? What can it say? What can it change? What requires authorization? When must a human take over?

This is why compliance cannot simply be a certification displayed on a vendor page.

It has to shape the architecture and behavior of the agent.

During the Q&A, Chris explained that Botco.ai treats LLM providers as subprocessors and has Business Associate Agreements in place with model providers when PHI is involved.

Rebecca added another critical dimension: organizations need to know where the AI should stop.

In a healthcare workflow, an agent may be appropriate for providing vetted information but not for independently making a medical diagnosis. At that boundary, the appropriate resolution may be escalation to a clinician.

The webinar demonstrated this principle through Botco.ai healthcare workflows, including an AI + human PHQ-9 screening workflow where concerns can be routed to the appropriate clinician with structured information.

7. Guardrails should determine behavior

Guardrails are often discussed as though they are simply restrictions placed around an AI model.

Enterprise guardrails need to be more operational than that.

They help define what information an agent can access, what actions it can take, what topics are within scope and when it should transfer responsibility to a person.

Rebecca described AI as sometimes behaving a little like a toddler: assumptions can be wrong, so teams need to provide specific do’s and don’ts.

That becomes particularly important in regulated environments.

AskGRACE, for example, provides oncology patients with evidence-based information verified by UCLA clinicians and includes escalation for serious symptoms and connection to the care team when appropriate.

The goal isn’t simply to make the AI capable of answering more questions.

It is to make it reliably understand which questions it should answer and which it shouldn’t.

8. A new model version doesn’t eliminate the need for QA

AI models are improving quickly, but upgrading the underlying model isn’t the end of the story.

Chris noted that newer general models can sometimes outperform older fine-tuned models. But organizations still need regression testing when models change.

Why?

Because a model that performs better overall can still behave differently on specific workflows or edge cases.

That means AI quality assurance isn’t a one-time implementation task. It is an ongoing discipline involving evaluation, testing and feedback.

Or, as Chris put it:

“You don’t just want to move to the model—you have to do some level of regression testing.” — Chris Maeda

9. Stop measuring AI engagement as success

Traditional digital metrics can become misleading when applied to AI agents.

A long conversation might indicate engagement.

It could also indicate that the customer couldn’t get what they needed.

Chris suggested looking beyond engagement toward metrics such as answer accuracy, whether the agent accomplished its goal, user satisfaction and the business objective the agent was designed to achieve.

Rebecca extended that thinking to specific outcomes: if an agent exists to schedule appointments, verify insurance eligibility or accomplish another task, measure whether that outcome actually occurred.

This changes the fundamental KPI:

Don’t ask only: “Did the customer interact with the AI?”

Ask:

“Did the customer accomplish what they came here to do?”

10. Treat AI more like a junior employee than an oracle

Chris closed the webinar with perhaps the simplest mental model of the entire session:

“I think of these models as, like, junior employees.” — Chris Maeda

A junior employee may be extremely capable, but you don’t give them unlimited authority on day one.

You provide context.

You define responsibilities.

You show examples.

You review their work.

You correct mistakes.

You expand their responsibilities as they demonstrate reliability.

Enterprise AI requires much the same discipline.

Rebecca’s closing recommendation was to resist the temptation to skip those stages:

“Don’t skip any of these important steps.” — Rebecca Clyde

She recommended pre-testing, limited or soft rollouts, gathering feedback from representative users and learning from actual interactions before deploying broadly.

Because production isn’t where your AI should begin learning what can go wrong.

The enterprise AI gut check

Before investing in a new agent—or deciding whether an existing deployment needs to be fixed—ask eight questions:

Is it connected to the right systems?
Can it take action?
Does it follow the actual workflow?
Are its permissions defined?
Are guardrails in place?
Does it have feedback loops?
Has it been failure-tested?
Are outcomes being measured?

Those are the eight checks Botco.ai presented at the end of the session as a practical test of enterprise readiness.

And they point to the larger lesson from the webinar:

The model is only one piece of the agent.

The real work is everything surrounding it: the systems it can reach, the information it receives, the workflow it follows, the permissions it operates under, the guardrails that constrain it, the feedback that improves it and the outcome it is ultimately responsible for delivering.

AI that answers isn’t necessarily AI that resolves.

And for enterprises investing in AI, resolution is where the value begins.