Business Performance Analyst Agent
Course demonstration · Capstone 2 · Agentic AI Workshop 2026

How we got here: project history

A course project, not a product: Capstone 2 of the Agentic AI Workshop 2026, which asks for a "Business Performance Analyst Agent" that answers managers' questions about sales data. The company, AdventureWorks, is Microsoft's public sample database of a made-up bicycle maker. No real business data is used.

The build happened on 1 October 2026, with final touches on the morning of 2 October. Times are UK time, taken from file timestamps and the git history, so they are approximate.

timeline
  title 1 to 2 October 2026, from afternoon to the next morning
  Start : Starter notebook on made-up data
  Afternoon : Cut the AI cost by 32% : Brief turned into 13 steps : Section 6 built in the notebook
  Early evening : Demo app around the agent : Cassette J-card design
  Evening : Real AI runs recorded, 6 of 7 : Clear for every audience : Checked against the brief, 7 of 7
  Night : Home screen, blog and API docs : Every track re-recorded live
  Next morning : Final touches and a clean-PC check

The starting point (15:15)

The course gave us a starter notebook and a training guide. The notebook already had a working agent "engine": a loop that lets an AI model pick tools, retry logic for the free Groq AI plan, a scorecard and a weekly briefing. It ran on made-up data: a pretend retailer with North, South, East and West regions, prices in pounds, and four planted "stories" to find.

The brief asked us to:

Step 1: Cut the AI cost (about 16:00 to 16:50)

The free Groq plan allows about 8,000 tokens a minute and 200,000 a day. We measured where the tokens went: every round of the agent loop re-sends the rulebook and the whole tool menu. Shortening both cut the fixed cost per round by 32% (about 2,650 to 1,810 tokens), and one anomaly scan of five metrics went from five rounds to one. Notes: prompt-optimization.md.

Step 2: Turn the brief into a plan (afternoon)

We wrote the brief out requirement by requirement (18 items, P1 must-have to P3 stretch) and turned it into 13 steps a non-programmer can follow: where to click, what to paste, why, and how to check it worked. The main decision: keep the engine and swap what's underneath. Sections 1 to 5 of the notebook stay as they were, and a new Section 6 goes at the bottom. Guide: build-guide.md.

Step 3: Build Section 6 in the notebook (afternoon, finished 20:49)

In a copy of the notebook, notebooks/capstone2_business_analyst_adventureworks.ipynb:

AW_PLANT = {"territory": "Northwest", "month": "2024-05", "category": "Bikes", "extra_discount": 0.25}

Step 4: A demo app around the same agent (from 19:48)

A notebook is hard to present. We moved the same code into a small Python package (app/analyst) behind a FastAPI server, with a hand-built web page and three modes: Live AI, Replay (plays back a recorded live run, step by step, as a backup if the wifi or AI fails) and Dry run (no AI, for testing the screens only). 27 automatic checks run without any AI key.

Step 5: Product brief and design (20:08 to 20:33)

We wrote down who it's for (judges and a "leadership team" watching a projector) and what success looks like (product.md). For the look we chose a hand-lettered cassette J-card instead of the usual dark AI dashboard: the six-minute demo as a mixtape, with side A for the brief's questions and side B for the guardrails (design.md).

Step 6: Real AI runs, recorded (20:36 to 20:58)

The free Groq plan ran out of its daily allowance during testing, so the agent now falls back to OpenAI automatically (gpt-5.4-mini). We ran every demo question live once and recorded it for Replay. On those first recordings the scorecard passed 6 of 7. The miss: asked for a return rate, the AI treated orders with negative quantities as "returns" and stated "0.00%", instead of saying the data has no returns (fixed in Step 8). We also made a 60-second pitch video with music and a voiceover from ElevenLabs.

Step 7: Make it clear for every audience (21:00 onwards)

Step 8: Check every line of the brief (evening)

We went through the printed brief line by line and recorded the evidence for each requirement in requirements.md. That found three gaps, now fixed:

Step 9: A finished product (night)

What we learned

Scope

A learning project, not a product. AdventureWorks is Microsoft's public sample data about a made-up bicycle company. Made by Victor Saly, Akashdeep Nijjar, Manuel Verduzco Valenzuela, Mazen Ahmed, Buddhika Gamage, Nathan Fryatt, Michael Kampouridis and Malak Sheat (the team). Source of this page: docs/history.md.