Title graphic: One problem, carried through every stage. Crawl, walk, run: a practical Copilot deployment, from Outlook to Copilot Studio agents.

If you’ve sat through a Copilot demo, you’ve probably seen the feature tour. Summarize an email. Make a slide. Ask a chatbot something. Twenty minutes in, everyone agrees it’s impressive and nobody can say what it would change about Tuesday.

We’ve given that demo. It doesn’t work, and the reason is simple. Features don’t stick. Problems do.

So when a pump and process-equipment distributor in the Southeast asked us to show their leadership what Copilot could do for them, we built the whole ninety minutes around one problem every distributor knows. A customer’s pump keeps eating mechanical seals. Three failures since July. The maintenance manager is angry and wants a plan by Monday.

We built it on data shaped like a distributor’s business: customers, orders, service tickets, manufacturer documents. That’s how we start every engagement. Not with a product, but with a problem a department already has and the person who owns it. AI is the first technology the business can shape around its own problems, big or small, common or niche. The people who get the most out of it aren’t in IT. They’re the reps, service techs and branch managers who already know the work, so they should drive it. IT’s job is the guardrails: security, compliance, governance and auditing. That split, business drives and IT guards, is the frame for this whole series.

The other half of the frame is layers. Copilot isn’t one thing. There’s Copilot in the apps, the agents Microsoft ships, agents you build in Agent Builder, and agents you build in Copilot Studio, and each one costs and does something different. Claude, ChatGPT and the rest have layers too. Most organizations pay for layers they don’t use and miss the ones that would solve their problem. Crawl, walk and run is how we match the problem to the right layer.

This is the first of seven posts on what we built, how we built it, and where it gives time back. It’s written for the people who decide what happens next: IT and business leaders, and anyone who has been asked to “show us what AI can do” and knows the trap in that sentence.

Where the structure came from

The demo didn’t start as a demo. Two months earlier we’d run an AI roadmap session with the same company’s IT and marketing leads. The deck for that meeting had a slide we’ve used in every engagement since: crawl, walk, run, with a gate between each phase.

Crawl means prove it’s safe. Name an owner, approve a policy, capture a baseline. Walk means prove it’s useful. Run use-case workshops, build a few agents, train by role. Run means prove it scales. Roll out in waves, fold governance into business as usual, report ROI at six months.

That session produced a short list of what the business wanted to see next. Notebooks. Analyst. Designer if time allowed. Our adoption scorecard and document intelligence work. Fair enough. But a list of five features is a feature tour waiting to happen.

The fix was to map the features onto the phases instead of presenting them as a list. Crawl became Copilot inside the apps everyone already has. Walk became the agents Microsoft ships, plus one built live. Run became agents built for a distributor’s business, in Copilot Studio, connected to data shaped like theirs. Same framework they’d seen in August, now with a working build attached to each stage.

Four steps: roadmap meeting, five requested topics, three stages, ninety-minute run of show

Why one story

Here is the problem with a feature tour. Each demo resets the audience’s context. You show Excel, they think about Excel. You show an agent, they think about agents. Nothing accumulates.

With one story, every step adds to what they already know. By the time the Copilot Studio agent pulls up the service history for pump P-101 and finds the same seal failing three times, the room has already seen the customer’s angry email summarized in Outlook, read the pre-visit notes drafted in Word, and watched Analyst flag the seal spend spike in the sales export. The agent isn’t a new thing. It’s the same thing, further along.

We picked a seal failure because it touches every function a distributor has. Sales sees the revenue. Service sees the tickets. Engineering sees the application problem. Marketing, eventually, sees a campaign. One chemical-plant customer and one pump gave us a thread through all of it.

The technical story underneath is real engineering. In the scenario, the customer changed their process fluid in June, from roughly 8 percent caustic to roughly 32 percent at about 150 degrees. The seal that was fine before isn’t rated for that. The right answer is a different seal with a different flush plan and a root-cause visit before the next turnaround. That answer lives in the manufacturer documents, and the whole build is about getting people to it faster.

That’s where the time comes back. Today, answering that maintenance manager means digging through a long email thread, pulling ticket history out of the ERP, exporting sales data and pivoting it, then hunting through manufacturer PDFs for the right seal spec. With Copilot set up properly, each of those becomes a question asked in plain English by the person who already knows what to ask. What used to be an afternoon of exporting and pivoting becomes one prompt. The branch rep doesn’t need to file a ticket with IT to get there.

Left: five disconnected feature demos. Right: one problem passing through five stages, each carrying context forward

What we actually built

The kit that came out of this is bigger than a demo usually is, and the next six posts go into each part. In short:

  • Realistic source material. Five manufacturer PDFs that cross-reference each other, a 34-page field service handbook with three contradictions buried in it, a seven-email customer thread, a messy 1,975-row sales export, a 485-row ERP extract across five tables, and a brand guide.
  • Scripted built-in Copilot scenarios for Outlook, Teams, Word, PowerPoint and Excel, each with a prompt list and an answer key.
  • Three declarative agents for Agent Builder, mirroring Notebooks, Analyst and Designer so we could answer the “can we make it ours” question.
  • Four Copilot Studio agents with embedded ERP data, Adaptive Cards, proactive Teams alerts, and strict rules about what they will and won’t say.
  • A test suite of 64 cases with regex grading, deploy scripts, a validator, a simulator and a morning-of preflight check.
  • A five-slide opener and a run document with timings and a cut order for when things run late.

It took five days and 43 commits. That number matters less than the order of work, which was the opposite of what you’d expect. We built the Copilot Studio agents and their test suite first, because they were the riskiest. The built-in app demos came last, because they’re the most forgiving.

The rule we’d give anyone doing this

Pick the problem before you pick the features, and make the problem one the room already has. Everything else follows. The features you show are the ones that move the problem forward. The ones that don’t, you cut. We cut Designer to five minutes and Notebooks to a maybe, and nobody missed them.

The second rule is almost as important. Make the data look like the business, so the people who run the business can see themselves in it. We gave the scenario a believable process change, realistic part numbers, and a service history with dates. Halfway through, someone asked whether we’d pulled the data from their ERP. That’s the reaction you want, because it means the sales and service people in the room have stopped watching a demo and started picturing their own Tuesday.

That’s also the point where IT’s role changes. It doesn’t shrink. IT decides who can reach which data, keeps the agent connections read-only, sets the rules that make an agent cite its source and refuse to guess, keeps the audit trail, and holds anything back until it passes the test suite. After that, the service manager and the branch reps who know the customers drive it.

That’s the CB5 standard: start with one business problem, let the business drive the solution on the right layer, and run it inside IT’s governance so the value lasts and it stays safe.

Seven rows, one line per post in this series

The rest of the series

  1. Crawl. What Copilot does in Outlook, Teams, Word, PowerPoint and Excel, what it can find in your files that people miss, and the honest line about licence tiers.
  2. Walk. Researcher, Analyst, the Notebook that refuses to quote a price, and the agent built live in under four minutes.
  3. Run. Four agents in Copilot Studio, the rules they all follow, and the one rule for choosing the tool.
  4. Testing. Sixty-four cases, a preflight that prints GO or NO-GO, and the thing that broke anyway.
  5. Training. Five capabilities, taught in order, tools last.
  6. What the room asked. It wasn’t about features.

Next: the crawl stage.

CB5 Solutions helps organizations understand the layers of AI, get the most from what they buy, and build an AI program that pays back inside their IT standards. We start every engagement with one business problem and a governance baseline. Let’s pick yours.