Affordable AI Consulting for Legacy Systems | 2026 Guide
- November 21
- 19 min
An Agent Discovery workshop is a structured session that turns a broad problem space into a short list of automation scenarios, then into one or a few detailed candidates ready for agent design, with explicit scoring, prioritization, and ownership.
A team finishes an Agent Discovery workshop. Fifteen scenario cards sit on the board. Everyone agrees the day was productive. Nobody picked a first place. The cards stay equal priority, the artifact lacks an owner, and design never starts. The workshop delivered activity without producing a useful outcome. Agent Discovery creates real value when the session ends with a narrow set of detailed candidates and a named initiative owner. Facilitation separates those two endings. This article explains how the discovery funnel creates value, how facilitation protects or breaks it, and where AI fits as a supporting tool inside that discipline.
Key Takeaways:
The output teams expect from Agent Discovery is a list of agent ideas. The output worth having is a ranked shortlist where one scenario has enough detail to start design.
Those are different things. The first takes a day to generate. The second takes that same day plus a disciplined funnel that most workshops never finish.

A well run workshop follows an inverted pyramid. Ideation is wide and non evaluative. Scoring makes candidates comparable. Prioritization cuts volume. Detailing goes deep only on what survives. Each stage discards enough that the next stage can apply full attention to what remains.
The sequence is where the value lives. Wide ideation without scoring produces options nobody can prioritize. Scoring without ideation narrows around the first idea that came to mind. Design without a detailed survivor builds on assumptions rather than on a specific scenario.
Structured Agent Discovery formats follow this logic: an inspiration phase to align on operational challenges, ideation against capability prompts, a scoring step across complexity and business value, prioritization to a single primary candidate, and then a detailed scenario the design phase can use. The brand of the guide matters less than whether the group runs all five stages in order. The SAP AppHaus Business AI Explore methodology is one well documented example of this structure.
Capability cards and fixed scenario fields look like restrictions. In practice they function as shared vocabulary. They push participants to describe what the scenario does, how it connects to existing systems, and why it matters at a business level.
An open whiteboard invites tool discussions. Structured prompts produce scenarios. The difference shows up at scoring time: formatted cards are comparable, whiteboard clusters typically are not.
When ideation runs fully, it surfaces scenarios nobody planned to write down. In one product discovery run, a non obvious candidate appeared from working through the prompts: terminology harmonization across project documents. The idea was to map synonyms and variant phrasings into a shared vocabulary before deploying an agent, so the agent operates on consistent concepts rather than on semantic noise spread across transcripts. That insight came from structured exploration, not from open brainstorming.
Not every scenario that survives scoring belongs in an agent design sprint. Prioritization produces a second useful output: problem class.
Some scenarios are fixed procedures with deterministic paths. Classical automation handles those. Some require reasoning plus real action in systems, which is the agentic pattern. Some need reasoning without system action, which may point toward analysis tools or assisted decision support rather than an agent.
Each call is useful. It routes the survivor to the right next track. Business process automation and agent design solve different shapes of problem. Discovery is where that distinction gets made before design investment starts.
Facilitation is the condition under which the funnel works. The facilitator explains each phase, gives examples, holds sequence, and keeps a named outcome in view. Without that, the group runs the activities and leaves without a decision ready artifact.
This happens more often than workshop guides acknowledge.
Workshops without a defined outcome drift. Participants fill cards because the session has blocks and the blocks have time boxes. Nobody agrees on what the final artifact should contain.
Early evaluation is the most common structural failure. When the group starts judging ideas during ideation, variety and volume collapse. Promising scenarios get dismissed before they are written down. Tool preferences fill the silence.
Prework reduces the worst of this. Participants who arrive with shared context around operational challenges spend less time explaining problems and more time generating scenarios worth scoring.
The hardest facilitation call is prioritization. Many organizations operate under a norm where every initiative carries equal urgency. Bring that norm into a discovery board and nothing gets selected.
A practical move: ask the business sponsor or senior participant to rank candidates with a single first position. One card can be first. Others form a documented shortlist. Only the first position receives full scenario detail that day. The rest remain on record without detail.
The question of what goes first is one a sponsor can answer in a minute. The question of whether everything is equally urgent is one they will avoid for an hour.
Vocabulary drift is a facilitation failure. A multi step approval workflow becomes “an intelligent agent.” A report generation script becomes “autonomous operations.” When problem class stays undefined, the group scores scenarios that are misrouted before design begins.
Keeping the procedure / agentic / reasoning distinction visible through scoring and prioritization prevents this. It also saves design time. A deterministic process that enters an agent design sprint eventually gets unwound back to the right solution, at higher cost.
A workshop can produce a strong primary scenario and still deliver nothing after the session. If no one leaves as the initiative owner with clear accountability and decision authority, the artifact waits for a follow up that never arrives.
Facilitation controls the session outcome. Ownership controls what happens after the session. Name who carries the primary survivor into design, and later into any PoC evaluation with real decision rights, before the closing block ends.
Skilled facilitation is less about energy in the room and more about defending the useful property of each stage.
During framing, the facilitator keeps attention on operational challenges and process pain, not on product demos or model names. During ideation, ranking is postponed deliberately. Capability prompts stay open questions. Clustering waits until volume exists.
The scarce resource here is patience. Sponsors want a winner by lunch. The method needs options first, judgment second.
Scoring produces comparable candidates when the scale is agreed before cards are rated. Useful dimensions include operational complexity, degree of system interaction, the proportion of rules to human judgment, and business value of the outcome.
A score without a written rationale teaches nothing. A rationale without a shared scale cannot be compared across cards from different participants. The facilitator walks one example card through the full scale before the group starts scoring independently.
Silent rewrites of criteria mid session destroy comparability. Stopping them is a facilitation responsibility.
Detailing applies only to survivors. Fifteen cards rated equally produces fifteen thin descriptions. One primary card with full as is / to be detail, named system touchpoints, and human decision rules gives design something concrete to build from.
|
Stage |
Benefit of the method |
How facilitation protects it |
How the session breaks |
|
Ideation |
Breadth and non obvious candidates |
Delay judgment; use capability prompts |
Critique during generation; tool discussions |
|
Scoring |
Comparable candidates |
Shared scale with written rationale |
Gut ranking without criteria |
|
Prioritization |
Focus for design effort |
Force a single first place |
Everything stays P1 |
|
Detailing |
Ready input for design |
Detail only the primary survivor |
Thin detail spread across all cards |
|
Handoff |
Continuity after the session |
Named initiative owner |
Artifact with no carrier |
Facilitators also face a cold start problem. Experience running Agent Discovery comes from running Agent Discovery. The first session still requires discipline. A detailed guide with worked examples of completed cards, and a co facilitator for the first run, both help. The method still depends on someone willing to hold sequence when the room wants shortcuts or early winners.
AI belongs in this analysis after the funnel and the failure modes are clear. Used inside a running method, a model reduces friction on the tedious steps. Used as a substitute for method, it reproduces the same failures at speed: volume without ranking, scores without rationale, polished text without a decision.

The steps that slow workshops down are not insight generation. They are making many cards comparable, writing rationales for each score, building prioritization matrices, and turning a shortlisted scenario into a structured detail pack.
A product owner who understands the product can work with a model that holds the method sequence. The model can explain what the next phase expects, show examples of completed artifacts, propose how to cluster related scenarios, draft scores with short comments for each dimension, sketch a prioritization view, and fill scenario detail templates based on what the product owner describes.
That combination produces a different quality of exploration. In practice, a small group of two or three people working with a shared transcript and a model configured around the method tends to generate more and narrower candidates than a solo session or an undirected group session of the same length.
The work is still demanding. The model proposes. The group reads, overrides where product context changes the answer, and owns the final call.
Model draft scores are directionally useful when the scale has been made explicit. They still require override. Local product knowledge, organizational constraints, and domain conventions the model has not seen will produce wrong scores on some cards in every session.
Accepting draft scores without review recreates priority collapse with formatted output. Overriding scores is the calibration step that makes the artifact reflect product reality rather than the model’s priors.
The same principle applies to ideation volume. Generating many agent candidate labels without running them through scoring and prioritization creates an unranked backlog at speed. Reviewed AI use means applying the same discipline to discovery artifacts that teams apply to generated code: read, verify, and own the output before it drives decisions.
One more practical observation: visual boards help people decide together in a room. Tables and markdown help models analyze and draft the next phase. Optimizing artifacts only for one audience imposes costs on the other. Facilitation serves both audiences simultaneously, which requires keeping both formats through the session.
Spec driven approaches help once an agent role is chosen. They structure requirements, task decomposition, and implementation steps. They consume a chosen role. They do not generate candidates worth choosing. BMAD is one example of this class of framework.
Starting in a build framework without prior discovery often commits design to the first scenario that sounded agentic. Discovery is what filters the pool before build investment begins. AI support for facilitation does not change that sequence.
Governance and accountability still ask who owns the decision when the artifact leaves the room. That question has no model answer.
Measure the workshop by what continues after participants close their laptops.
A useful handoff includes five elements:
Anything thinner leaves design guessing. A wider set usually means prioritization did not finish.
Capture the group decision in a format the group can see and discuss. Export a clean textual version for later design use and for any AI assisted drafting in the next phase. Different audiences need different forms of the same information.
A detailed scenario without an owner is a document waiting to be archived. An owner without a detailed scenario is a sponsor with no brief to carry.
Both elements belong in the final block of the session, before the room disperses and attention moves elsewhere.
When one primary scenario is clear, detailed, and owned, the team can move into agent design with a narrower scope and a real decision path. Any proof of concept that follows then has a baseline, a named outcome owner, and defined success criteria rather than a hopeful demo.
The quality of Agent Discovery is the quality of the narrowing it produces and the continuity of the person who carries it forward.
An Agent Discovery workshop is a structured session that moves from a broad problem space to a short list of automation scenarios, then to one or a few detailed candidates for agent design. It uses ideation, scoring, and hard prioritization so design effort goes to the scenarios that survive, rather than to every idea that sounded agentic in the room.
A good facilitator protects sequence and outcome. They explain each phase, give worked examples, hold ideation open long enough, enforce a shared scoring scale, and force a single first place on the board. They also keep problem class visible: procedure automation, agentic work, and reasoning without action each point to different next steps and different design investments.
Move when one primary scenario is detailed enough to design behavior, system access, and human oversight, and a named initiative owner is in place. Discovery can still document other candidates on the shortlist. Design spend should follow the primary survivor through the ownership gate, then into validation with defined success criteria for any proof of concept.
AI can support facilitation on clustering, draft scores with comments, matrix proposals, and template fill when a product knowledgeable human still calibrates judgment. It can also help someone learning the method stay on sequence through phases. It becomes counterproductive when scores are accepted without override, volume replaces prioritization, or the model draft is treated as the sponsor decision.
They fail when facilitation drops the funnel. Early judgment kills ideation breadth. Everything stays first priority. Fixed procedures get labeled as agents. The session ends with many thin cards and no detailed scenario. Or the artifact has no initiative owner, so design and validation never start even after a productive day.