← AI Catalyst C3All programsHomeSearch
AI Catalyst C3·Core Session - Week 13·1:08:00

Community Session 1: Causa Claims — a Non-Coder's Five-Agent Construction-Claims Pipeline, the Gate as the Product, and What RAG Really Costs

Syed Hasan Presenter - civil engineer, 30 years in construction project controls (Middle East, UK, US), 'totally a non-coder'; built Causa Claims, a five-agent extension-of-time claims pipeline that won the Outskill x OpenAI hackathon (~10,000 applicants, 5 winners) · Shivani Host - explains the community-session format, fields chat, calls for member teachers (LangChain/LangGraph suggested next) · Praful Participant - runs the AIC Mastermind call; asks the RAG cost and dynamic-ingestion questions; proposes self-hosting vector DBs on a VPS

The short version

  1. Extension-of-time claims are won on records, dates, causation and entitlement - not prose - and a contractor lost a multi-million claim by filing notice on day 31 of a 28-day window. Syed's answer is five agents (Contract Reader, Evidence Scout, Chronology Builder, Drafter, Quantum) behind deterministic, code-enforced gates: 'the agents are not the product... the gate is the product.'
  2. Live on a real UAE bridge project: FIDIC clauses 20.2.1 / 8.1 / 20.2.4 pulled, top-12 evidence ranked by RAG, a Haiku-built chronology, a citation-backed draft that correctly finds the claim TIME-BARRED (day 38 vs 28), and a Quantum agent (Emden/Hudson formulas) showing ~$2M entitlement on a hypothetical valid claim - every step logged with model, latency, cost and a hashed trust log for the judges.
  3. Praful's cost question: he didn't track tokens; ballpark a ~$100/week Codex subscription for the seven-day build, Supabase ~$10/project plus ~$25 of API calls; production needs daily/weekly or event-triggered re-ingestion; Supabase chosen for speed, Pinecone considered, Qdrant on a VPS is Praful's cheaper route.
  4. Lessons: 'the fluency was precisely the risk' - confident prose earns trust it hasn't paid for; Generic AI vs Domain AI vs Product AI; 'AI capability alone is not a product - the value is the workflow, plus the verification, plus a human at the right point.' He'd redo it smaller - one agent, refusal gate first; a real claims consultant found three gaps in ten minutes and he is rebuilding for launch 'within a month or two'.
  5. Chained with LangChain in Python by prompting step by step ('loop engineering'), not OpenClaw/Hermes or Codex's agent runtime; most of the time went into checking outputs, not learning AI.

At a glance, three clicks deep

Skim here first: the closed row is the glance, open is the study card with the key points and timestamps, and the ↓ link drops to that concept's full write-up below.

01Five agents behind a deterministic gate: the gate is the productLLM agents propose;

LLM agents propose; code-enforced gates (citations, date logic, eligibility) decide; every fact carries provenance.

Claims decided on records, dates, causation, entitlement (0:10)

Day-31 notice on a 28-day window lost a valid claim (0:12)

Five agents; 'the gate is the product' (0:13)

Evidence cards with provenance and confidence (0:16)

Chronology rule enforced in code (0:17)

↓ Full write-up of this concept

02The live run: FIDIC clauses, top-12 evidence, a time-barred verdict, and a $2M what-ifPer-step model routing with visible cost/latency, provenance on every fact, a hashed audit log;

Per-step model routing with visible cost/latency, provenance on every fact, a hashed audit log; the pipeline can refuse.

Contract Reader: FIDIC 20.2.1 / 8.1 / 20.2.4 (0:38)

Evidence Scout top-12 via RAG (0:40)

Haiku chronology; Opus for drafting (0:42, 0:46)

Time-barred verdict: day 38 vs 28 (0:44)

Quantum: Emden/Hudson, ~$2M on a valid claim; trust log (0:46-0:50)

↓ Full write-up of this concept

03What the RAG actually cost, and why ingestion must recurHackathon RAG: ~$100/wk Codex + ~$10 Supabase + ~$25 API;

Hackathon RAG: ~$100/wk Codex + ~$10 Supabase + ~$25 API; recurring ingestion required; VPS self-hosting cuts per-instance cost.

~$100/week Codex subscription; tokens untracked (0:52-0:53)

Supabase ~$10/project + ~$25 API; VPS/Qdrant alternative (0:54-0:57)

Praful: ~$30 for his own ingestion (0:55)

Scheduled or event-triggered re-ingestion in production (0:56-0:57)

Supabase for speed; Pinecone considered (0:58)

↓ Full write-up of this concept

04Generic AI, Domain AI, Product AI - and why fluency was the riskMoat = domain rules encoded as gates;

Moat = domain rules encoded as gates; ship one verified agent before five; expert review finds gaps fast.

'The fluency was precisely the risk' (0:17, 0:59)

Workflow + verification + human = the product (0:59)

Generic / Domain / Product AI (1:01)

Redo: one agent, refusal gate first (1:02)

Consultant found 3 gaps in 10 minutes; rebuild (1:02-1:03)

↓ Full write-up of this concept

05How a non-coder chained five agents: LangChain, prompted step by stepPrompt-driven LangChain scripts in a fixed domain order;

Prompt-driven LangChain scripts in a fixed domain order; verification is where the hours go; LangGraph next.

'Loop engineering' via prompting, LangChain in Python (1:07)

Not OpenClaw/Hermes/Codex agents (1:08)

Time spent on checking outputs (1:08-1:09)

LangChain/LangGraph proposed as next community topic (1:07, 1:10)

↓ Full write-up of this concept

The concepts in full

01

Five agents behind a deterministic gate: the gate is the product

Every line in a claim can be challenged. So no line gets in without a citation and a date that makes sense.

Contract Reader (which clauses apply, notice compliance), Evidence Scout (every fact as an auditable 'evidence card': document, date, page/clause, confidence), Chronology Builder (hard rule enforced in code, not by the model: an event cannot cause a delay that started before it), Drafter (citation-backed narrative), Quantum (cost). The model proposes; the system verifies. Motivation: claims are decided on records and causation, and opposing lawyers hunt for unprovable sentences; the day-31 notice story shows the failure is a handover problem, not an AI one.

Why it matters

The clearest statement in the programme of where to put determinism when an LLM writes something a professional will attack.

02

The live run: FIDIC clauses, top-12 evidence, a time-barred verdict, and a $2M what-if

The correct answer was 'no extension' - and the demo said so.

Real event: late-issued M&E/coordination drawings. Contract Reader pulls FIDIC 20.2.1, 8.1, 20.2.4 and flags notice compliance; Evidence Scout retrieves and ranks 12 items (PDF ingestion, embeddings, similarity search on pgvector); Chronology Builder on Haiku sequences the timeline with citations; the Drafter concludes the contractor is time-barred (submitted day 38 against 28); the Quantum agent applies Emden and Hudson formulas and, re-run with a hypothetical 40-day claim, shows about $2M. Each agent step displays model, latency and cost; a trust log (row hashes, JSON audit trail) was built for the hackathon judges. Opus (as heard 'Opus 4.7') for the methodology-selection reasoning, Haiku for the cheap step, an OpenAI model for ranking. Primavera P6 programme data was an input.

Why it matters

Model routing by step cost - the cheap model for chronology, the expensive one for judgement - is the pattern to copy.

03

What the RAG actually cost, and why ingestion must recur

He built it on a hundred-dollar-a-week subscription and never counted a token.

Praful (who has shipped similar ingestion pipelines - ~$30 just to ingest) asks for the real numbers. Syed: no granular tracking; the seven-day build ran on a ~$100/week OpenAI Codex subscription (required by the hackathon); Supabase as backend and vector store at roughly $10 per project plus ~$25 in API calls. Praful's alternative: self-host Supabase or Qdrant on a VPS and multiply instances per fixed cost. In production, documents must be re-ingested on a schedule (daily/weekly) or on events so retrieval stays current; Supabase was chosen for speed under deadline, Pinecone considered, no deep vendor comparison.

Why it matters

Concrete numbers for the 'what will this cost to run' question every client asks.

04

Generic AI, Domain AI, Product AI - and why fluency was the risk

The better the prose sounded, the more dangerous it was.

Biggest failure mode: hallucination wrapped in fluency - reviewers trust confident text more, not less. 'AI capability alone is not a product... the value is the workflow, plus the verification, plus a human at the right point.' Three layers: Generic AI (what the cohort learned - 'every one of us has the same model access'), Domain AI (30 years of claims expertise), Product AI (what the programme teaches you to build). If redoing the hackathon: one agent done well, the refusal/fluency gate first, scale later. A practising claims consultant found three gaps in ten minutes; he is rebuilding from the ground up for launch within a month or two.

Why it matters

Paul's edge is domain knowledge of small-business operations, not model access - this names that explicitly.

05

How a non-coder chained five agents: LangChain, prompted step by step

No agent framework, no orchestration code he wrote by hand - a roadmap and a lot of checking.

Orchestration was done by prompting ('loop engineering', his term) into standalone Python/LangChain scripts, following his own domain roadmap (contract -> narrative -> cost), not OpenClaw, Hermes or Codex's native agents. He recommends the community learn LangChain and LangGraph for chaining and suggests Outskill teach them; Shivani takes it as the next member-taught topic. Most build time went to iteratively verifying outputs rather than learning new techniques. Side threads: Rahul's unresolved question on lip-synced Indian-language AI video (Praful has a Hindi example in the AIC Mastermind group); community-session format may merge with Outskill's 'Forge' meetings; Cody's Web-MCP is in a separate OpenAI contest.

Why it matters

Evidence that the programme's loop-engineering vocabulary is now how members describe their own builds.

Tools referenced

ToolCoverageMomentContext
LangChainexplainedRAG pipeline and agent chaining in Python
SupabaseexplainedBackend + pgvector store, ~$10/project
LangGraphmentionedRecommended next topic
Codexmentioned~$100/week subscription, hackathon requirement
PineconementionedConsidered, not used
QdrantmentionedPraful's self-host suggestion
ClaudementionedOpus for drafting, Haiku for chronology
Primavera P6mentionedSchedule data input

Session materials

Archived locally on V: — click to open. Companion pages link to the LMS.

Action items

    Resources mentioned

    Resources
    • docCausa Claims (heard CostaClaims.com)

    Extraction notes

    This page was built from an auto-generated transcript, which garbles product and people's names. Those were corrected silently in everything above and logged here for transparency. The warnings flag claims that were true on the recording day but change fast.

    Transcript corrections applied

    The transcript saysThe trainer actually means
    Kausa Claims / CostaClaims.comCausa Claims
    Prime EveraPrimavera P6
    Clodus / Cloud Opus 4.7Claude Opus (version unverified)
    PG vector / PD vectorpgvector
    quadrantQdrant
    civilian engineercivil engineer

    True on recording day — verify before relying