The session's instrument is also its lesson: one desktop app where chat, files, browser, terminal and git live together — pick your reasoning depth per task and build.
Codex is OpenAI's counterpart to Claude Code, and the tour covered what a non-engineer needs: projects are just folders (create fresh or point at an existing one); the model is GPT-5.5 with a reasoning selector — low/medium for simple work, high for coding, extra-high for gnarly debugging — which doubles as a cost dial. The top-right panel holds the working surfaces: file browser with previews, an in-app browser (the feature he says Claude Code lacks — preview the site, scroll it, screenshot sections), code-change review, and a terminal. Every project gets git tracking automatically; the branch/changes vocabulary gets its own session later.
His tooling politics are refreshingly provisional: he downgraded Claude Code, runs Codex at $100 (a $20 ChatGPT plan is enough to start), is moving his team over — and openly says he may switch back when Anthropic ships something better. The meta-lesson: agents are interchangeable workbenches; skills and workflows (this session's real content) travel between them.
The live setup: 'outskill elegant landing pages' project created from scratch, dummy files added, file previews opened, git changes visible — the whole workbench in ninety seconds.
Everything from here to the end of the course happens inside an agent workbench like this one. Knowing the surfaces — and the reasoning dial — is the difference between driving the tool and being driven.
Pick the one true coding agent and commit.
Even the trainer's choice is provisional ('ask me in two weeks'). Skills and workflows are tool-agnostic — the workbench is a rental, the systems are yours.
The reasoning-tier dial is the same discipline as your Fable-vs-Opus routing: match model depth to task stakes. And his in-app-browser praise explains why your preview-file pattern matters — seeing the artifact beats imagining it.
Go deeper
In one line: OpenAI's Codex app (chatgpt.com/codex, works with any ChatGPT plan): projects are folders, GPT-5.5 with selectable reasoning (low/medium/high/extra-high), a top-right panel with files, in-app browser, code-change review and terminal — and automatic git tracking on every project.
His stack decision: downgraded Claude Code, moved to the $100 Codex tier, team migrating next month — 'and I may have to go back' (0:09:38)
Reasoning tiers as a cost dial: low/medium for simple tasks, high for coding, extra-high for hard debugging (0:13:44)
The in-app browser is the differentiator he cites over Claude Code: preview, scroll, screenshot the site being built (2:04:19)
One interface for everything: chat, projects, automations (weekly invoice reconciliation shown), plugins/apps — Slack, GitHub, Notion, Gmail (1:14:55)
▶ Watch this taught: 0:09:38
Answer from memory first — the recall attempt is what makes it stick. Then reveal.
How do you choose a reasoning tier?
Low/medium for simple tasks, high for most coding, extra-high for complex debugging — it's a quality-vs-cost dial.
What four surfaces live in the top-right panel?
Files (with preview), the in-app browser, code-change review, and the terminal.







