← AI Sprints (Live Weekend Programs)All programsHomeSearch
AI Sprints (Live Weekend Programs)·Working With AI·2:42:25

AI Sprint: Choosing the Right Problem — Day 2 (A Nine-Step Worksheet from Pain Inventory to PRD, and the Honest Screen for Whether AI Belongs at All)

Dileep (KVSS Dileep) Head of Generative AI Education at Outskill - teaches the whole W1-W9 worksheet live, runs three timed pen-and-paper exercises, and argues most AI side projects die of bad problem selection, not bad tools · Sumida Sprint host - opens with the roll call and recordings walk-through, then runs the certificate demo and the paid-programme pitch from 2:08 · Learner cohort (chat) Contribute example problems, poll answers and questions throughout; three first names surface during certificate troubleshooting (Kartikeya, Venkataraman, Amulya - ASR spellings uncertain)

The short version

  1. Day 1 was how to delegate to a bot. Day 2 is the prior question nobody asks: is this problem worth solving with AI at all? The answer is a nine-step worksheet, W1 to W9, that carries ONE problem from a raw list of annoyances to a build-ready brief - and W1 through W3 are 'completely non AI. It is pure problem solving. W4 onwards is where AI fitment comes into the place.'
  2. The evidence for why this matters is two cited studies: a CSCW hackathon study that tracked ~590 projects and watched activity fall from 35% right after the event to 17% within a week to 3.5% within months, and an MIT 2025 report in which 95% of 300 surveyed AI projects showed weak P&L impact. The diagnosis in both cases is problem selection, not tooling.
  3. W1 is a pain inventory - about ten recurring annoyances from the last fortnight, written by hand, with AI explicitly forbidden ('your brain is the LLM'), each quantified by frequency, minutes lost and who else it hurts. Painkillers beat vitamins. W2 narrows to one, and forces you to notice when you have written a solution ('remind my team') and called it a problem.
  4. W3 is a Socratic five-whys chain to the root cause, where every 'because' must connect to the one before it or the chain is invalid. The worked examples land somewhere unexpected on purpose: 'I read all 34 notes by hand every day' bottoms out at 'the summary was never the rep's job. It was mine', and a calendar problem turns out to be a delegation problem.
  5. W4 is the AI-fitness screen and the honest part: Part A disqualifiers, Part B qualifiers, Part C 'what would a rule, a filter or a template do here, and is that genuinely not enough?', the automate-versus-augment call, Part E error-and-ROI arithmetic, and Part F where you paste your own root cause into a thinking model and tell it to argue against you and 'do not soften it'. Dileep calls that shaking the nail.
  6. W5 to W9 converge again: a solution-agnostic 'how might we', eight possible shapes the answer could take, a buildability gate set at a realistic 4-6 focused hours, a five-beat storyboard (trigger, input, transformation, output, human moment) and finally the brief - 'somewhat like a PRD'. Underneath it all sits Version 0, the no-AI way you do the job today, which whatever you build has to beat.

At a glance, three clicks deep

Skim here first: the closed row is the glance, open is the study card with the key points and timestamps, and the ↓ link drops to that concept's full write-up below.

01The W1-W9 worksheet: one problem, nine steps, no AI until step fourW1 pain inventory - W2 pick one - W3 five whys - W4 AI-fitness screen - W5 how might we - W6 eight shapes -…

W1 pain inventory - W2 pick one - W3 five whys - W4 AI-fitness screen - W5 how might we - W6 eight shapes - W7 buildability - W8 storyboard - W9 brief.

One worksheet, one problem, one week; QR code on every slide (0:25)

'W1, W2, W3 is completely non AI... W4 onwards is where AI fitment comes into the place' (0:26)

Steps are sequential - skipping ahead undermines the later ones (0:26)

Blended and simplified from IDEO design thinking and Google PAIR (1:20)

End state is a filled brief: problem, goal, users, inputs, outputs, data/access, success criteria, why-not-a-rule, out of scope (0:26)

↓ Full write-up of this concept

02Why AI side projects die: the hackathon curve and the 95% reportProject mortality is a problem-selection failure, not a tooling failure;

Project mortality is a problem-selection failure, not a tooling failure; screen the problem before you build.

CSCW study, ~2020, ~590 hackathon projects tracked (0:29)

35% active immediately, 17% by six days, 3.5% by five months (0:29-0:30)

MIT 2025 report: 95% of 300 AI projects had weak P&L impact (0:32-0:33)

Speaker hedges the MIT figures with 'I think' - treat as recalled, not verified (0:32)

Conclusion: motivation and fit decide survival, not the model you chose (0:33-0:34)

↓ Full write-up of this concept

03The bar for 'solved with AI' is a saved prompt, not a productSolution ladder: saved prompt / custom GPT or project / skill / spreadsheet AI / automation / vibe-coded pr…

Solution ladder: saved prompt / custom GPT or project / skill / spreadsheet AI / automation / vibe-coded product; seven-day box; dummy data only.

Six acceptable solution forms, lowest rung a saved prompt (0:37-0:40)

Seven-day time box, justified by Parkinson's Law (0:41)

No passwords, no PII, dummy data (0:44)

Ask 'how might I solve this and where could AI fit', not 'can I use AI here' (0:45)

You must be able to state the problem, the outcome and the check before submitting (0:46)

↓ Full write-up of this concept

04W1 pain inventory: ten annoyances, by hand, quantified - painkillers not vitaminsTen recent recurring pains, handwritten, each with frequency, time lost and blast radius;

Ten recent recurring pains, handwritten, each with frequency, time lost and blast radius; keep painkillers, drop vitamins.

~10 pains from the last two working weeks (0:53)

Seed prompts: what you'd teach a new coworker first; what you retype because two systems don't talk (0:56)

Quantify each: what, how often, time lost, who else (0:56)

AI banned - 'LLM... equal to prefrontal cortex' (0:59)

Painkillers over vitamins: vitamin products 'rarely succeeded' (1:04)

↓ Full write-up of this concept

05Intuition atrophy: judgment is a muscle you can delegate awayProlonged delegation of judgment to a model degrades the delegator's own judgment;

Prolonged delegation of judgment to a model degrades the delegator's own judgment; keep some decisions manual on purpose.

Heard on an entrepreneurs' podcast; term garbled by ASR, most likely 'intuition atrophy' (0:50)

'Delegated your judgment... once you stop using LLMs, your intuition reduces' (0:51)

Compared to brain atrophy and to an unused muscle (0:51)

Practical response: periodically take decisions back from the model (0:52)

Explains why W1-W3 are handwritten and AI-free (0:26, 0:59)

↓ Full write-up of this concept

06W2: 'remind my team' is a solution wearing a problem's coatA problem statement describes the friction and its cost;

A problem statement describes the friction and its cost; if it names an action or a tool, it is a solution and must be reopened.

Five minutes to pick exactly one problem for the week (1:08)

'Remind my team' is a solution, not a problem (1:09-1:10)

Attendance case study: dedicated app failed; WhatsApp selfie at entry/exit worked (1:11-1:12)

Prefer embedding into an existing habit over introducing a new app (1:12)

Disguised solutions foreclose the root-cause work in W3 (1:10)

↓ Full write-up of this concept

07W3: five whys, Socratic - and every 'because' must connect to the last oneAsk why ~5 times;

Ask why ~5 times; each answer must follow from the one above; stop only at a cause specific enough to act on.

Socratic framing; ask why about five times (1:15-1:16)

CRM example: 34 notes by hand -> 'the summary was never the rep's job. It was mine' (1:17)

Broken chain = incomplete framing, start again (1:27)

Calendar overload turns out to be a delegation problem (1:26)

'The system is bad' and 'how the org is structured' are too vague to stop on (1:16)

↓ Full write-up of this concept

08W4 Parts A-C: disqualifiers, qualifiers, and 'would a rule have done it?'Part A disqualify, Part B qualify, Part C prove a rule or template is insufficient - before AI is allowed i…

Part A disqualify, Part B qualify, Part C prove a rule or template is insufficient - before AI is allowed into the design.

Part A disqualifiers: narrow or drop the task (1:30)

Part B qualifiers: what makes a task AI-shaped (1:30)

Part C: name a rule, filter or template and ask if it is really not enough (1:30)

Catastrophic single mistakes are an explicit disqualifier (1:37)

'Do not try to lie about your problem statement' - sometimes the answer is no AI (1:39)

↓ Full write-up of this concept

09Automate or augment: unpleasantness and agreement decideAutomate = tool performs, human reviews the result.

Automate = tool performs, human reviews the result. Augment = human decides at each step, tool assists. Choose on unpleasantness x agreement.

Automate: tool performs, human reviews only the final result (1:31)

Augment: human keeps deciding at every step (1:31)

Two deciding questions: is it unpleasant, do people agree what correct looks like (1:31)

Directional is good enough - 'you don't have to get it right completely' (1:31)

↓ Full write-up of this concept

10W4 Part E: name the wrong answer, then time both sidesDefine wrong and missed outputs, decide which is worse, then measure manual time versus checking time;

Define wrong and missed outputs, decide which is worse, then measure manual time versus checking time; the delta is the ROI.

Define wrong and missed outputs before building (1:32)

Decide which failure is worse in this workflow (1:32)

Time the manual baseline and the checking effort (1:32)

'Reclaiming 10 minutes... that is a great victory' (1:32)

Turns a vague sense of usefulness into a defensible case (1:32)

↓ Full write-up of this concept

11W4 Part F: make the model argue against you, and 'do not soften it'Adversarial review: model argues the strongest case against your framing;

Adversarial review: model argues the strongest case against your framing; you rebut each objection in one line or accept the hole.

Paste the W3 root cause plus six disqualifier criteria into a thinking model (1:33)

Instruction is explicit: 'do not soften it' (1:33)

Write a one-line rebuttal to each objection raised (1:33)

'Shaking the nail' - test that the conclusion holds when challenged (1:35)

Parts A and B alone are acceptable if time runs out (1:33)

↓ Full write-up of this concept

12Version 0: the no-AI way you do it today, and the bar the build has to clearVersion 0 = the documented manual process, built first, used both as the comparison baseline and as the sou…

Version 0 = the documented manual process, built first, used both as the comparison baseline and as the source of real examples.

'Version 0 is a version without AI at all' (0:43)

'Whatever you build has to kind of be beaten by the tool' (0:43)

Scheduled for days 1-2 of the seven-day week, before AI (2:05)

Produces 3-4 real examples that later feed the spec and the tests (2:05)

If AI does not beat it on time, quality or insight, do not build (0:43)

↓ Full write-up of this concept

13W5: a 'how might we' that names no toolA solution-agnostic reframing of the root cause that admits several possible answers and is legible to an o…

A solution-agnostic reframing of the root cause that admits several possible answers and is legible to an outsider.

Rewrite the W2-W3 root cause as 'how might we...' (1:49)

'Should allow different types of solutions' (1:51)

Hints where to start but 'names no solution tool or technology' (1:51)

Must 'make sense to someone outside the team' (1:51)

Worked example: attend every meeting -> fewer meetings require my presence (1:50)

↓ Full write-up of this concept

14W6: eight shapes the answer could take, before you pick oneDiverge to eight candidate output forms, then converge on one by build effort, one-week feasibility and dat…

Diverge to eight candidate output forms, then converge on one by build effort, one-week feasibility and data availability.

Sketch eight possible output forms - spreadsheet row, chat message, draft email, one-pager, checklist (1:52)

Divergent and non-committal: 'my solution looks like this', not an app spec (1:54)

Pick one on ease of build, one-week feasibility, available data (1:54)

Shapes are droppable - WhatsApp automation traded for email when integration proved hard (1:54)

Human thinking is still the default tool here; AI may assist from this point (1:55)

↓ Full write-up of this concept

15W7: four to six focused hours, ten archetypes, and the signs it is too bigGate on 4-6 focused hours;

Gate on 4-6 focused hours; map to an archetype; abort on the five oversizing signs.

Real budget is 4-6 focused hours, not a week - 'life keeps happening on the side' (1:56)

Archetypes: read-many, fill-a-table, sort-and-route, notes-to-actions, recurring report, Q&A over docs, compare lists, checklist, form-and-record, morning digest (1:57)

Too big if it needs more than one new system connection (1:58)

Too big if you cannot name the test that proves it works, or get 10-20 examples by day two (1:58)

Too big if it is an open-ended assistant, or stakes are high and mistakes hard to spot (1:58)

↓ Full write-up of this concept

16W8-W9: five beats, then the brief that is 'somewhat like a PRD'W8 = trigger / input / transformation / output / human moment.

W8 = trigger / input / transformation / output / human moment. W9 = the brief: problem, goal, users, inputs, outputs, data and access, success criteria, why-not-a-rule, out of scope, human check.

Five beats: trigger, input, transformation, output, human moment (1:59)

Same five beats for an automation, an app, a prompt or a skill (2:00)

W9 brief fields include 'why not a rule' and 'out of scope' (2:02)

'This brief is what is somewhat like a PRD' (2:05)

A complete W9 signals 'complete problem understanding' - ready to build (2:02)

↓ Full write-up of this concept

The concepts in full

01

The W1-W9 worksheet: one problem, nine steps, no AI until step four

how-to

Nine boxes on one sheet. The first three of them forbid you from opening a chat window.

The whole session is one artifact: a worksheet that carries a single problem from raw observation to a build-ready brief across nine numbered steps. The split is the point - 'w 1, w 2, w 3 is completely non AI. It is pure problem solving. W 4 onwards is where AI fitment comes into the place.' Steps are sequential and skipping ahead breaks the later ones, because W4 screens the root cause found in W3 and W9 writes up the shape chosen in W6. One worksheet per problem per week; take a fresh copy for the next one. The framework is not invented - it blends and simplifies IDEO design thinking and Google PAIR.

Do it in this order
Why it matters

This is a reusable client-facing intake instrument: it turns 'we should use AI for something' into a documented, argued, scoped brief before anyone spends money.

Choosing the right problem: nine steps, and AI is not allowed for the first three NO AI - PURE PROBLEM SOLVING W1 pain inventory ~10 pains, quantified W2 pick ONE not a disguised solution W3 five whys chain must connect W4 AI-FITNESS SCREEN A / B / C - automate or augment - E - F the honest gate: some problems die here AI MAY ASSIST FROM HERE W5 how might we names no tool W6 eight shapes pick one W7 buildable? 4-6 focused hrs W8/W9 storyboard + brief = PRD VERSION 0 - how you do it by hand today built days 1-2; whatever you build has to beat it Seven-day box: Version 0 and 3-4 real examples first, then interview the owner, turn the brief into a spec, add AI to the smallest step, test with the real user. W1-W3 by hand - W4 decides whether AI belongs at all - W5-W9 shape and specify it
The nine steps, with the non-AI block, the AI-fitness gate and the Version 0 baseline that any build has to beat.
W1, W2, W3 is completely non AI. It is pure problem solving. W4 onwards is where AI fitment comes into the place.
02

Why AI side projects die: the hackathon curve and the 95% report

590 hackathon projects. Six days later, 17% still had a heartbeat. Five months later, 3.5%.

Dileep opens with evidence rather than exhortation. The first citation is a CSCW conference study from around 2020 - 'even before AI came with the picture' - that tracked roughly 590 hackathon projects: about 35% showed some activity straight after the event, 17% by day six, 3.5% by five months. The second is the MIT 2025 report that went viral, in which 95% of 300 surveyed AI projects showed weak P&L impact. His reading of both is the same: the failure is upstream of the tooling. People pick a problem they are not motivated by, or one that was never AI-shaped, and the build dies quietly. Hence a whole session on selection.

Why it matters

The two numbers are the argument for charging a client for discovery instead of jumping to a build.

03

The bar for 'solved with AI' is a saved prompt, not a product

A saved prompt counts. A spreadsheet formula with AI in it counts. You do not have to ship an app.

Before setting the exercise, Dileep widens the definition so nobody opts out for lack of engineering. A solution can be a saved prompt, a custom GPT or project, a skill, spreadsheet AI, a workflow automation, or a vibe-coded product - six rungs on the same ladder. He then time-boxes the whole thing to seven days, invoking Parkinson's Law, and sets the data rules: no passwords, no personally identifiable information, use dummy data. Finally he reframes the guiding question. Not 'can I use AI here?' - which invites you to force it - but 'how might I solve this, and where could AI fit?'

Why it matters

The reframe is the difference between a client asking for an AI project and a client getting a result.

People get this wrong

If it is not an app or an agent, it does not count as an AI solution.

A prompt you saved and reuse, or a formula, is a solution if it beats the way you do the job today.

04

W1 pain inventory: ten annoyances, by hand, quantified - painkillers not vitamins

how-to

'Your brain is the LLM.' Close the laptop for ten minutes.

W1 asks for about ten recurring pains from roughly the last two working weeks - recency bias is deliberate, though the window can be widened. The seed prompts are designed to surface friction rather than wishes: 'if I were training a new coworker, what would I teach them first?' and what gets retyped or re-checked 'because two systems do not talk to each other'. Each entry must be quantified - what concretely happened, how often, minutes lost each time, who else is affected - because W4 will later do ROI arithmetic on exactly those numbers. AI is banned for this step outright. And the filter at the end is painkiller versus vitamin: things that would merely make you happier 'rarely succeeded'.

Do it in this order
Why it matters

Quantifying at capture time is what makes the later business case possible; most discovery skips it and never recovers.

05

Intuition atrophy: judgment is a muscle you can delegate away

'Once you stop using LLMs, your intuition reduces.'

Mid-exercise, Dileep pauses on a risk he heard described on an entrepreneurs' podcast. The transcript mangles the term - 'intuition thrust', 'intuition trust', 'intuition rushed' - and he apologises for his own pronunciation, but the idea he describes is unmistakable: 'if you have delegated your judgment to LLMs for a very long time, once you stop using LLMs, your intuition reduces.' He compares it directly to brain atrophy and to a muscle that weakens from disuse. It is offered as a caution inside a pro-AI framework, not against it, and it is the reason W1 to W3 are handwritten: the exercise deliberately claws some decisions back.

Why it matters

A useful, honest counterweight for a client who wants to automate judgment as well as labour.

If you have delegated your judgment to LLMs for a very long time, once you stop using LLMs, your intuition reduces.
06

W2: 'remind my team' is a solution wearing a problem's coat

If your problem statement contains a verb you could implement, it is not a problem statement.

W2 is a five-minute narrowing to exactly one problem for the week, and the trap it exists to catch is the disguised solution. 'Remind my team' is not a problem - it is a fix someone has already chosen, which quietly forecloses every other option and skips the root cause entirely. The supporting case study is the attendance app: a purpose-built app failed, and the working answer turned out to be a selfie taken through WhatsApp's camera at the building entrance and exit. The lesson is to embed the fix in a habit that already exists rather than asking people to adopt a new one - a constraint that should shape the shape you pick at W6.

Why it matters

The most common failure in a client brief, and the cheapest one to catch.

07

W3: five whys, Socratic - and every 'because' must connect to the last one

how-to

'Socrates was the most annoying teacher... whenever the disciple asks a question, Socrates always counter questions and asks why.'

W3 is twelve minutes of driving one pain down to a root cause by asking why about five times. The rule that makes it work is structural: the chain of 'because' statements has to hold together end to end, and 'if a because does not connect to the previous one, the Socratic framing is incomplete' and has to be redone. The worked examples all land somewhere the surface never suggested. 'I read all 34 notes by hand every day' bottoms out at 'the summary was never the rep's job. It was mine' - an ownership problem. A calendar-overload complaint bottoms out as a delegation problem, not a scheduling one. Dileep explicitly rejects stopping at vague causes like 'the system is bad' or 'how the org is structured'; those are not specific enough to build against.

Do it in this order
Why it matters

Solving the symptom is how you build something correct that nobody uses.

The summary was never the rep's job. It was mine.
08

W4 Parts A-C: disqualifiers, qualifiers, and 'would a rule have done it?'

'Part A is your disqualifiers. Part B is your qualifiers. Part C is a solution where you use absolutely no AI at all.'

W4 is the honest part, and it is where most candidate problems should die. Part A lists disqualifiers - conditions under which the task should be narrowed or dropped, one of which is named explicitly: a single mistake being catastrophic. Part B lists the qualifiers that make a task a good fit. Part C is the one people skip: name the non-AI alternative first - a rule, a filter, a template - and then ask honestly whether it is genuinely not enough. The whole screen only works if you are willing to lose: 'do not try to lie about your problem statement... if you are honest, then you kind of come to the conclusion that AI should not be used' in some cases.

Why it matters

The rule-filter-template question is the single cheapest piece of consulting there is, and it is free.

People get this wrong

The screen is a formality on the way to building the AI thing.

A correctly run screen kills some problems outright; that is the outcome it exists to produce.

W4, the AI-fitness screen: three gates, one call, one sum, one argument against you PART A disqualifiers narrow it or drop it; one catastrophic mistake = out PART B qualifiers what makes this task AI-shaped PART C would a RULE do it? name the filter or template first - is it genuinely not enough? AUTOMATE or AUGMENT unpleasant + people agree what correct looks like -> automate. Otherwise augment. PART E - the arithmetic name the wrong answer and the missed item; time the task by hand vs time to check one AI output PART F - shake the nail paste the root cause in, tell the model to argue against you and "do not soften it" - rebut in one line "Do not try to lie about your problem statement. If you are honest, then you kind of come to the conclusion that AI should not be used." Parts A and B alone are an acceptable stopping point when time runs out
The AI-fitness screen: three gates, then the automate-or-augment call, the ROI arithmetic and the adversarial check.
Do not try to lie about your problem statement. If you are honest, then you kind of come to the conclusion that AI should not be used.
09

Automate or augment: unpleasantness and agreement decide

'Automate tasks that are difficult, unpleasant, where people agree what correct looks like. Augment tasks that people enjoy or where people disagree about what correct means.'

Once a task clears the screen, one call remains: does AI do the task end to end with a human reviewing only the result (automate), or does it assist while the human keeps deciding at every step (augment)? Two questions settle it - is the task unpleasant, and do people agree on what a correct outcome looks like? Where both hold, automate. Where the work is enjoyable, or where reasonable people would disagree about the right answer, augment. Dileep is explicit that this does not need to be exact: 'you don't have to get it right completely. At least directionally.'

Why it matters

Two questions, said out loud in a kickoff, prevent the most expensive category of AI disappointment.

Automate tasks that are difficult, unpleasant, where people agree what correct looks like. Augment tasks that people enjoy or where people disagree about what correct means.
Check yourself

Answer from memory first — the recall attempt is what makes it stick. Then reveal.

Copy-editing a colleague's article - automate or augment?

Augment. People disagree about what a correct edit is, and many writers enjoy the control. Automate the mechanical checks only.

10

W4 Part E: name the wrong answer, then time both sides

how-to

'What does the wrong answer look like? What does a missed item look like? What is worse in this workflow?'

Part E forces two numbers into existence before any building starts. First, define failure concretely: what a wrong output looks like, what a missed item looks like, and which of the two is worse in this particular workflow - because that determines whether you tune for precision or recall. Second, time both sides: how long the task takes by hand today, and how long it takes a person to check one AI-generated output. The gap between those two is the entire ROI case. Dileep's framing of the payoff is deliberately modest: 'every time the task is done, every time I'm reclaiming 10 minutes, and that is a great victory.'

Do it in this order
Why it matters

Checking time is the cost everyone forgets; without it, the savings are imaginary.

11

W4 Part F: make the model argue against you, and 'do not soften it'

You have just spent an hour proving your problem is AI-shaped. Now pay a model to demolish it.

The last part of the fitness screen turns your own reasoning against itself. Paste the W3 root cause and the six disqualifier criteria into a capable thinking model and instruct it to argue, paragraph by paragraph, that AI is a bad fit for this problem - with the explicit instruction 'do not soften it'. Then write a one-line rebuttal to each objection. Anything you cannot rebut in one line is a real hole. Dileep names the move by analogy: you do not hammer a nail in and walk away, you shake it to see whether it holds. He also concedes the full A-to-F sequence is hard and tells learners it is fine to get only through Parts A and B inside the time box.

Why it matters

The cheapest possible red team, and the only step in the worksheet that reliably changes minds.

12

Version 0: the no-AI way you do it today, and the bar the build has to clear

'Version 0 is a version without AI at all... whatever you build has to kind of be beaten by the tool.'

Version 0 is the manual way the task is done right now, written down properly, and it is the yardstick everything else is measured against. It is not a thought experiment: it is scheduled as literal day-one and day-two work in the coming week, before AI is allowed anywhere near the problem. Doing it by hand has a second payoff - it produces three or four real worked examples of the task, and those examples become the test set and the specification later. If the AI version does not beat Version 0 on time, quality or insight, there is nothing to build.

Why it matters

Without a baseline there is no way to tell whether the AI version is an improvement or just newer.

Version 0 is a version without AI at all... whatever you build has to kind of be beaten by the tool.
Check yourself

Answer from memory first — the recall attempt is what makes it stick. Then reveal.

Why build the manual version when you already know AI can do it faster?

Because 'faster' is unmeasurable without a baseline, and because doing it by hand three or four times is how you discover the edge cases the spec would otherwise miss.

13

W5: a 'how might we' that names no tool

'How might we set up a system so that you don't have to attend so many meetings?'

W5 rewrites the root cause from W3 as a solution-agnostic 'how might we' statement. The discipline is in what it must not contain: no tool, no technology, no named fix. Three quality tests are given. It should allow different types of solutions rather than pointing at one. It should hint at where to start while naming no solution, tool or technology. And it should make sense to someone outside the team - if it needs insider context to parse, it is written wrong. The worked example follows the calendar chain: a root cause of being expected at every meeting becomes 'how might we set up a system so that fewer meetings require my presence?'

Why it matters

Written properly, this one sentence is what you show a client to agree scope before anyone argues about tools.

14

W6: eight shapes the answer could take, before you pick one

A row appearing in a spreadsheet. A chat message. A draft email. A weekly one-pager. A checklist.

W6 is deliberately divergent: sketch eight different forms the answer could take, Post-it style, committing to none. The examples given are all small - 'a sheet with a row appearing in a spreadsheet, a chat message, draft email, weekly one-pager, a checklist' - because the point is to break the reflex that every solution is an app. 'You are just saying my solution to the problem looks like this.' Only after the spread exists do you choose one, on three criteria: ease of build, one-week feasibility, and whether the data is actually available. Shapes stay droppable: a WhatsApp-automation idea gets traded for email once the integration turns out to be hard.

Why it matters

Most briefs arrive having already chosen the shape, usually the most expensive one.

15

W7: four to six focused hours, ten archetypes, and the signs it is too big

The week is not a week. 'Life keeps happening on the side.' It is four to six focused hours.

W7 is a feasibility gate with an honest clock. The nominal window is one week; the real budget is four to six focused hours, and everything else must be scoped to fit. To help, Dileep offers a menu of solution archetypes to map the idea onto - read-many, fill-a-table, sort-and-route, notes-to-actions, a recurring report, question-and-answer over documents, comparing lists, a checklist, form-and-record, a morning digest. Then the warning signs that a problem is oversized: needing more than one new system connection; being unable to name a test that would prove it works; being unable to gather ten to twenty examples by day two; being an open-ended assistant rather than a fixed repeatable job; or high stakes where mistakes would be hard to spot.

Why it matters

The archetype list doubles as a menu Paul can show a client - most real requests are one of these ten.

16

W8-W9: five beats, then the brief that is 'somewhat like a PRD'

how-to

'Till this point is the hard part. After this, it is about solutioning.'

W8 storyboards the chosen shape as five beats, and they are the same five whether the build ends up an automation, an app, a prompt or a skill: 'what is the trigger, what is the input, what is the transformation, the output, and the human moment, who checks, approves, or gets flagged.' W9 then writes the brief - problem, goal, users, inputs, outputs, data and access, success criteria, why not a rule, out of scope - plus the human check: what must be produced, when it passes, what a human approves. Dileep equates it to a PRD and treats a fully completed W9 as the evidence of 'complete problem understanding'. The remaining week layers on top: interview the problem owner if it is not your own problem, collect examples, turn the brief into a spec, add AI, test with the real user, submit.

Do it in this order
Why it matters

This is the deliverable - a one-page brief a builder or a vibe-coding tool can act on without a meeting.

Till this point is the hard part. After this, it is about solutioning.

Tools referenced

ToolCoverageMomentContext
CertFlowdemonstratedCertificate download tool; ASR renders it 'search flow'
WhatsAppexplainedAttendance case study: selfie through the existing camera beat a purpose-built app
LM StudiomentionedSuggested for running a downloaded session transcript through a local model if you missed the live session
ChatGPTmentionedNamed as an example of 'LLM' - and explicitly excluded from W1
ClaudementionedSame - named, then banned for the pain-inventory step
GeminimentionedSame - named, then banned for the pain-inventory step

Action items

Resources mentioned

Resources
  • docThe W1-W9 problem-selection worksheet
  • docSeven-day build plan
  • docIDEO design thinking and Google PAIR
  • docRecordings, transcripts and summaries
  • docCertificates via CertFlow
  • docPaid-programme pitch (2:08-2:42)

Extraction notes

This page was built from an auto-generated transcript, which garbles product and people's names. Those were corrected silently in everything above and logged here for transparency. The warnings flag claims that were true on the recording day but change fast.

Transcript corrections applied

The transcript saysThe trainer actually means
Dilip / Dalit / The lipDileep (KVSS Dileep)
party is the disqualifiersPart A is the disqualifiers
search flow / cert flowCertFlow
intuition thrust / intuition trust / intuition rushedintuition atrophy (speaker flags his own pronunciation)
high at the rate Outskill dot com / high android Outskill dot comhi@outskill.com
c s, CWCSCW (ACM Conference on Computer-Supported Cooperative Work)
Uratechunclear robotics company, possibly Unitree - unverified
Amalia / Amulyaone chat participant, spelling uncertain

True on recording day — verify before relying