Antigravity finally stopped being a chore. /teamwork-preview has a coordinator split the job across specialist agents, /boost checks everything before it acts, and Pro/Ultra quotas went up (check your own account for the numbers). It's crossed from toy to something that can carry light daily work.
A field guide · 2026 edition
From prompt to agent,
a clear path.
Two years of notes on using AI, gathered for friends drowning in information who haven't found the door yet — one no-detour path, plus the tools, thoughts, and pitfalls along the way.
Start the trailNo matching notes — try another word.
Field notes
Field notes · scroll sideways · share to X
Live voice can stay on while you work, but it's a separate model (GPT-Live-1), not the main one answering. I use it to hand out tasks and keep a running work log, not to finish hard jobs in one shot. And Desktop Work voice draws on its own quota — don't open it unless it's actually driving the computer.
Codex sessions can hand work to each other now: one passed its deploy plan to another, which picked it up and paused its own release to avoid a collision. Running several tasks at once, I'm hoping to stop being the one relaying messages between them.
First lesson from GPT-6: watch the tokens. Ultra and multi-agent runs aren't my default — a few agents going at once burns through budget fast. Start on a lower reasoning tier and step up only when the task earns it. Size the effort to the job.
Two weeks into Claude Code, the biggest realization: people who write tests instantly pulled ahead of people who don't. Once the agent finishes, you have exactly one job left — verify. Verification runs on tests.
Prompt engineering isn't about magic words, it's about stating the task clearly. People who can write a good prompt for Claude get better at product specs, emails, and technical docs too.
Local models aren't about saving money — they're about privacy and determinism. Ollama plus a 70B Q4 model handles 60% of daily questions; only the other 40% needs the cloud.
The beginner's biggest mistake: treating AI as a search engine. It isn't. It's a biased, forgetful, but fluent colleague. Learning to work with a colleague like that is the core skill of this era.
Gemini 2.5 Pro's 2M context — I dropped a whole 800-page contract in and asked, and it genuinely remembered a clause on page 30. Five years ago this was science fiction.
MCP is underrated as a protocol. Once agents can call any tool in a standardized way, the moat at the model layer gets hollowed out further — what's left is the moat of 'tools + data'.
A few words first
Plain-English basics · explained like you're five
Token
The smallest unit of text an AI reads and writes. Like snipping a sentence into little pieces (a word, half a word) — the AI "thinks" one piece at a time, and each piece is a token.
LLMLarge Language Model
A program that has read enormous amounts of text and learned to "guess the next word." You start it off, it keeps writing. GPT, Claude, and Gemini are all LLMs.
API
The "socket" that lets two pieces of software talk. Your program sends a request in an agreed format; the other side sends the result back in kind. Plugging into an AI means calling its API.
MCPModel Context Protocol
A "universal plug" standard for connecting AI to tools. With it, an AI can reach your calendar, files, and databases the same way — no separate connector built for each one.
Skills
A pre-written "how to do this one thing" manual + materials for the AI. When the task comes up it opens the matching skill and knows the steps and standards. (This page was built by an AI following a skill.)
Agent
An AI that breaks work into steps, calls tools, and carries a whole task through on its own — not just replying, but actually getting the job done.
loopagent loop
The rhythm of AI at work: think a step → reach for a tool → look at the result → think the next step, circling until the job's done. That loop is an agent's "heartbeat."
trace
The full running record of what an AI did step by step — what it thought, which tool it called, what it got back. When something breaks, you retrace it, like a dashcam.
embedding
Turns a piece of text into a string of number-coordinates; the closer in meaning, the closer the coordinates. AI uses it to judge how alike two passages are, and to pull the most relevant bits out of a big pile of material.
RAGRetrieval-Augmented Generation
Before answering, the AI pulls relevant passages from your own library and answers holding those — so it speaks from your material instead of making things up from memory (uses the embedding above to find them).
prompt engineering
Figuring out how to say things to an AI clearly, so it gives you what you want. Same task, different phrasing, wildly different results — that "how to ask" craft is this.
context engineering
One step past "how to ask": managing the whole packet of information you feed the AI each time — what to include, in what order, how much, what to cut. Most of whether it answers well comes down to this.
harness engineering
Building a whole "shell" around the model — tools, the loop above, permissions, how context is managed — turning a model that only chats into an agent that actually does the work. That shell is the harness.
TTSText-to-Speech
Reads written text out loud in a near-human voice. Audiobooks, turn-by-turn navigation, Siri speaking — that's it.
STTSpeech-to-Text
The reverse — turns what you say into text. Dictating to your phone, auto meeting transcripts, voice input — all of it.
GitHub
The site where programmers store code, edit together, and see every change's history. Like Google Docs + cloud storage + revision history, for code.
HTMLHyperText Markup Language
The "skeleton" language for building web pages — it says what headings, paragraphs, images, and buttons a page has. This very page is HTML.
Start here · the beginner's path
The six steps · click a node to expand
01
Understand what an LLM actually does
Mental model
It isn't looking up answers on Google for you. It's more like someone who's very good at continuing a sentence — you start it off, it guesses "what word most likely comes next," and keeps going. Get this, and the rest stops being confusing.
02
Learn to write a prompt
Prompt basics
Four things are enough: ① the context ② who it should act as ③ the format you want ④ an example. Nail those four before chasing tricks.
03
Pick one main model
Pick one
Don't shop around. Pick one (GPT or Claude or Gemini), use it honestly for three straight months, then talk about switching.
04
Wire it into your workflow
Integrate
Put it on real work: writing email, editing docs, reading PDFs, reviewing code. Make it a tool you reach for daily, not an occasional toy.
05
Ship a tiny project
Ship one
Build something very small — under 100 lines is fine. Finishing one with your own hands jumps your understanding a whole level.
06
Move up to agents / tooling
Agent layer
Once the first five feel easy, go advanced: let AI do the work itself (agents), connect tools (MCP), run models on your own machine. Don't rush this floor before the foundation is solid.
Claude · GPT · Gemini · Grok · DeepSeek
Pick one. Use it for three months. Then reassess.
Claude Opus 5.5My pick
Out Sept 22 · my main pick · Fable 5.1-level work at 40% less than Opus 5 to run
Strengths
- Anthropic's own numbers: Fable 5.1-level results on most work, for 40% less compute than Opus 5
- Built up for coding, agents, computer use and knowledge work — and it writes output 30% faster than Opus 5
- Plainer answers: they reworked how it communicates, so the point comes first and the jargon comes out
- Writing feel and code reasoning are still the signature — Claude Code is the killer app, and it's the setup I run every day
Weakness
- Sensitive cybersecurity and biology requests get routed to Opus 4.8 by the safeguards
- Native image generation still trails GPT; Sonnet 5.5 and Haiku 5.5 haven't shipped yet
Best for · Writing, code, long documents, agents — make it the daily driver, and save Fable for the problem that's genuinely hard
API $4 / $20 per million tokens (in/out), 20% below Opus 5 · on Pro, Max, Team and Enterprise plansGPT-6 family
Astra on top · Sol the workhorse · Luna the fast, cheap one · and 6.1 Sol landed Sept 29
Strengths
- Astra (rolling out since Sept 3): OpenAI's strongest model — a step up in coding, science and computer use
- Sol (Sept 22) for everyday coding and agent work; Luna for high-volume routine jobs like summarizing and extraction — both at half the price of the previous generation
- GPT-6.1 Sol (DevDay, Sept 29): OpenAI says it comes close to Astra on coding, computer use and professional work, at a fifth of Astra's price
- ChatGPT Images 2.5 shipped with Astra — image generation and iterative edits are still where GPT leads
Weakness
- Sol, Luna and 6.1 Sol live in ChatGPT Work and Codex for now, not the regular ChatGPT chat
- Astra's new reasoning hides its own reasoning trace — harder to audit
Best for · Astra for the hardest coding, science and computer-use work; Sol for daily agents and code; Luna for bulk routine jobs
API: Sol $2 / $10, Luna $0.10 / $0.50 per million tokens · Plus and up get them in Work and Codex; Luna is on the desktop app for Free and GoGemini 3.8 Flash
Built for speed · coding and agents are the focus · Google ecosystem
Strengths
- The newest, smartest "workhorse" model yet, aimed squarely at coding and agents
- Upgraded reasoning core, now with agentic video understanding
- Launched cheaper — the intro price is half the prior 3.6 Flash
- Wired into Gmail / Drive / Docs, plus a free AI Studio to try it and grab an API key
Weakness
- Ships fast (3.6 → 3.7 → 3.8 within weeks) — easy for notes like these to go stale
- Instruction-following occasionally loose
Best for · Speed, giant documents, Google users, devs who want AI Studio for free
Free (incl. AI Studio) + paid tiersGrok 4.7
Out Sept 21 · xAI's new flagship · 500K context · same price as 4.6
Strengths
- Trained for hard problems that run for hours: bigger base model, longer RL run, stronger self-checking, better long-context handling
- Reportedly about 2.1 trillion parameters, 40% more than 4.6
- Four reasoning-effort levels: low / medium / high (default) / xhigh
- Already in the Grok app, Grok Build, Cursor and GitHub Copilot
Weakness
- Text and image in, text out — nothing else
- Several reviews still put it behind Claude and GPT-6, and the launch slipped more than once
Best for · Long-running agent work, coding and knowledge tasks where you also want it cheaper than peers
API $2 / $6 per million under 200K prompt tokens, $4 / $12 above (in/out) · SuperGrok from $30/moDeepSeek V4
Open-weight · the app itself is free · general + reasoning merged into one
Strengths
- V4 merges the general-purpose (V-series) and reasoning (R-series) lineages into a single model
- Ships as Pro and Flash, open weights, MIT-licensed — you can self-host either
- 1M-token context at a very low cost
- The official app and web chat are completely free, no Plus/Pro paywall
Weakness
- API pricing now splits peak/off-peak, and peak-hour rates roughly double
- Self-hosting the Pro build is hardware-hungry (1.6T total parameters)
Best for · Free everyday chat, or self-hosting an open model to cut costs
App/web totally free · API off-peak $0.007/$0.22/$0.66 per million tokens, 5M free tokens for new accountsAI coding tools
Codex · Claude Code · Antigravity · Grok Build
Codex (OpenAI)
agentic Cloud / CLIOpenAI's official coding agent. Runs in the cloud — give it a repo and a task, it opens a PR itself.
Verdict · Good for: batched, well-defined changes, writing tests. Not for: exploratory work where you think as you go.
Claude Code
agentic TerminalAnthropic's terminal coding agent. Runs in your local shell — reads files, runs commands, runs tests.
Verdict · The strongest local coding agent right now — if you're willing to work in the terminal.
Antigravity
ide IDEGoogle's agent-native IDE — multiple agents work tasks in parallel, with a built-in verification flow.
Verdict · New, worth watching. The idea is "IDE for agents," not "IDE with AI."
Grok Build
agentic TerminalxAI's coding agent. Runs in the terminal, up to 8 parallel agents, working a plan → search → build loop.
Verdict · Local-first — your code never touches xAI's servers. Arena Mode auto-scores and ranks candidate solutions before you even look.
Other agent frameworks
Hermes Agent — An open-source autonomous agent from Nous Research — self-hosted, runs on your own server 24/7 against whatever model you point it at. Finishes a hard task, writes itself a reusable skill, and gets sharper the more you use it.
DeepSeek Harness — DeepSeek's open-source agent runtime, built "everything is a plugin" — model, tools, and interface are all swappable. Its own definition of an agent is Model + Harness — the exact split behind the "harness engineering" entry above.
Grok Bot — xAI's always-on AI teammates — each bot gets its own cloud computer, signs into your existing tools, and keeps grinding through multi-step work even after you've stepped away.
Gemini Spark — Google's 24/7 personal agent. Runs on Gemini 3.5 and is built on the same Antigravity harness listed under Coding tools above, wired straight into Gmail / Docs / Slides.
Git / collaboration terms
Git — A tool that snapshots your code and records every change, so you can jump back to any old version anytime.
repo — Repository: a folder holding all of a project's code + its history.
commit — Saving a checkpoint: recording what you changed this time with a one-line note. Like a save point in a game.
branch — A side road off the main line to change things, merged back once it works.
fork — Copy someone's whole project under your own name — change it freely without touching the original.
PR — Pull Request: the "I've made changes, please look and decide whether to merge into the main line" request.
Prompt library
Copy, fork, replace placeholders
You're my writing partner. I have a rough idea: "{IDEA}".
Please:
1. Distill its core thesis (one sentence)
2. Give three possible angles, each argued in a paragraph
3. Once I pick an angle, give a 5-part outline Below is code I just wrote. Play a senior engineer:
1. Point out the 3 most worthwhile issues (no more than 3)
2. Give fixes, noting the trade-offs
3. Don't chase "perfect" — chase "maintainable"
```
{CODE}
``` Explain "{TOPIC}" to me Feynman-style:
1. First in a way my 12-year-old could follow
2. Then deepen it with an everyday analogy
3. Finally, point out one common misconception I'm stuck on a decision: "{DECISION}".
Analyze it with the 10/10/10 framework:
1. Will I regret it 10 minutes from now
2. 10 months from now
3. 10 years from now
Two sentences max each. No pep talk. Projects
Small things, shipped — not planned
Choose-an-Agent Skill
A vendor-neutral Claude skill: teaches Ontario buyers/sellers how to judge, interview, and background-check a real-estate agent — and never recommends a specific one (me included). Open source, cloneable.
Claude Skill · TRESA/RECO · MIT
Email auto-sorter Agent
A small local agent that scans my inbox daily, sorts mail into 5 buckets, and labels it. Saved me about 4 hours in two weeks.
Claude API · Python · Gmail API
Local PDF Q&A tool
Feeds all my contract PDFs into a local vector store for offline Q&A. No internet — a lawyer friend gave it 5/5.
Ollama · Llama 3.1 70B · LangChain
Weekly-report generator
Reads git commits and Slack channels, auto-drafts a report every Friday. I only edit 20%.
Claude Code · GitHub API · Slack
Home-built MCP server
Exposes my home Synology NAS to Claude so it can read my family-photo metadata and organize albums.
TypeScript · MCP SDK
Going deeper with agents
Cloud · Local LLM · Frameworks · Evals
Cloud services
- Vercel AI SDK — the fastest way to plug into OpenAI / Anthropic / Gemini
- Replicate / Fal — pay-per-call backends for open-source models
- Modal / Beam — give your agent an execution environment in the cloud
Local LLMs
- Ollama — run a local model in one command, best for beginners
- LM Studio — a GUI, for people who dislike the terminal
- vLLM — self-hosted inference server, highest throughput
- Hardware bar: M-series Mac 32GB+, or a 24GB GPU
Agent frameworks
- MCP (Model Context Protocol) — Anthropic's tool-calling standard
- LangGraph — state-machine agents, high controllability
- Mastra — TypeScript-first, great DX
Evals / monitoring
- Langfuse — open-source LLM observability
- Braintrust — eval sets + experiment tracking
- Write the tests first, then the agent. Not the other way around.
Resources
No affiliates · no fluff
Free apps worth knowing
Real-estate tasks you can hand to an agent
What a realtor can hand to an AI agent
Full auto = the agent runs it, you look at the result; AI drafts · you review = the agent writes, you go line by line — it goes out under your name; Prep automated · you execute = the agent assembles everything, but the signature, the price and the send button stay with you. Not because the machine can't reach them — because that step shouldn't be handed over.
Listing copy
房源文案
Give it the specs and selling points, get listing descriptions, social captions, and email blasts (EN + ZH).
How · Claude / GPT on your copy template, run in batches.
Marketing design
营销设计
Listing posters, social graphics, just-sold cards and video covers, rendered in bulk from your brand template — several versions to pick from.
How · An image model plus a locked brand template (colours, fonts, credentials line), run in batches.
Bilingual comms
双语沟通
Translate and polish client messages, emails, and documents between English and Chinese, keeping tone and professionalism.
How · Any large model; use Gemini / Claude long-context for big files.
Neighbourhood briefs
社区资料
Compile schools, amenities, commute, and sale history into a one-page client brief.
How · A web-search agent on a uniform template.
Data entry
CRM 录入
Structure contact info from business cards, forms, and emails into your CRM, and tag it.
How · Agent + CRM API / Zapier, on a schedule.
Lead nurture
Lead 跟进
Draft first replies and nurture sequences, personalized by client stage.
How · CRM trigger + LLM draft, you glance before sending.
Multi-platform promotion
多平台宣传
One piece of content reworked for each platform — Xiaohongshu, WeChat, Instagram, YouTube — with the posting schedule laid out.
How · The agent drafts and schedules; you give it a look and confirm the send.
Pricing / CMA draft
定价草稿
Pull comparables into a first-pass CMA draft. You must check the numbers and the comp choices.
How · Data feed + LLM summary, human sets the price.
Scheduling
预约协调
Read intent from email, propose slots, write to the calendar; conflicts go to you.
How · Calendar MCP + agent.
Contract review
合同条款初审
Key clauses pulled out, watch-points listed, every clause lined up against the template — all assembled by the agent. Whether to change or sign is your call and your lawyer's.
How · Long-context model plus a clause library builds the comparison; a human decides clause by clause.
Compliance checklist
合规核对
The checklist and supporting documents for the RECO / OREA / FINTRAC process, assembled with gaps flagged up front. The final tick is yours.
How · The agent prepares the list and the paperwork; you tick and you sign.
AI is a biased, forgetful, but fluent colleague — learning to work with a colleague like that is the core skill of this era.
— Arthur · 2026.09.29
12 years full-time in real estate · Broker · FRI · SRS · ABR · MCNE · CLHMS & GUILD Elite · REAIS