The unusual part
I ran product at a talent-assessment company, so when I received my own psychometric profile I treated it as an engineering spec. I score 95 on Accountability and 13 on Establishing Order: I feel responsible for everything and naturally sequence nothing. So every agent in my system forces external capture on loose talk: owner, deadline, next action. Each score tension became a system requirement. The assessment is not a report I read once. It is the operating system.
Start here
The Machine
The full essay. One operator, seventy-three unattended jobs, roughly ninety agents: what I built, what it taught me, and what broke along the way. Seven chapters, about fifteen minutes.
Case studies
The Assessment That Became an Operating System
How psychometric score tensions became the design constraints for an agent fleet, and why the pure-LLM first version had to die.
Every Agent Wakes Up Knowing What the Others Did
Cross-agent persistent memory: the architecture, and the five failures that taught me memory is a distributed-systems problem in a productivity costume.
One Sentence In, One Company Out
One logged sentence to a shipping MVP: overnight research, then a 29-issue product program run by agents through two phase gates, then React and Vite code deployed on Vercel.
Cheap Models, Expensive Mistakes
A 20-trial harness measuring where local models can replace frontier ones. The local agent did not hallucinate, it silently did nothing 20% of the time, and that shape of failure decided the routing policy.
I Built the Product Org I Used to Hire
A 48-agent product organization with a Chief Product Officer agent. It ran a 29-issue gated product program for PsychBill end to end, and the findings became committed code the next day. What it revealed: craft specifies, judgment does not.
Demos
This renders from a script rather than a screen recording, so it regenerates against the current state of the machine instead of aging into a claim I can no longer reproduce.
Platform health: nine check modules over the whole machine in about six seconds. Uncurated, including two bugs the capture exposed in the reporter itself.
Systems running right now
The full architecture map lives at github.com/JustinTSmith/systems. The short version:
Memory spine
Every AI session starts knowing what previous sessions learned. An overnight job briefs each agent on what the others did.
Product Compass
A 48-agent product organization: a CPO agent routing to eight VPs and Directors, with roughly forty specialists beneath them.
Agent fleet
One Telegram thread into a three-agent fleet, plus 13 agent companies on Paperclip running roughly ninety agents.
Voice pipeline
One press of the iPhone Action Button: transcribed, classified 8 ways, filed into the vault in seconds.
Research to audio
Four parallel weekly research scans become a written briefing and a podcast-style audio overview I listen to instead of read.
Self-healing layer
Watchdogs restart dead services, nightly security review, health checks that commit their own reports.
Agent companies
13 companies and roughly 90 agents on Paperclip: PsychBill, a medical council, an equity research desk, staffed professional teams.
Local and cloud routing
Ollama and MLX serve local models for mechanical work; judgment escalates to frontier models. One pipeline analyzes on a local 35B and writes on Opus.
Shipped and open
- ai-operator-skills - six Claude Code skills distilled from pipelines I actually run
- Life-Operating-System - identity-driven daily planning with deterministic scoring and energy-gated modes
- personal-crm - relationship pipeline from Gmail and Calendar with a learning loop
- airline-checkin-skill - automated check-in across 8 airlines
- weekly-briefings - the weekly intelligence briefings, published as they generate
- qwen3-tts - local text-to-speech server, OpenAI-compatible API
AI and ML in production, at employers
- Two ML models shipped to production at Officevibe, inferring employee sentiment from latent behavioural data in a category built on surveys; directed a PM embedded in the internal AI lab building next-best-action recommendations
- MLOps process established at Workhuman for moving research models into production; highest-impact model shipped was adverse psychological bias detection in user messaging
- AI bidding engine at Acquisio: shaped the requirements from discovery and owned how the capability reached market
- 12 AI startups led at Creative Destruction Lab, working with Mila, the Quebec AI institute founded by Yoshua Bengio
Product track record
- 11% to 31% signup-to-activation in 4 months - Workhuman
- +50% PQL conversion and +15 NPS - Officevibe, while pivoting the product and retaining roughly 100% of clients
- $650K bookings before full launch, 9-month relaunch - SuccessFinder
- Enterprise to SMB repositioning supporting the sale of Acquisio to Web.com
- $30M+ equity value contributed across cohort companies - Creative Destruction Lab
Skills
- Product - strategy, product-led growth, OKR design, empowered team models, go-to-market, market repositioning, SDLC, hiring and team building
- AI and ML product, in production - model deployment, MLOps process design, model business-impact evaluation, customer-facing ML design, latent-signal and behavioural inference, next-best-action systems, AI requirements from discovery, AI positioning and GTM
- Applied AI, hands-on build - multi-agent orchestration, model routing and cost engineering, agent evaluation harnesses, local model serving (Ollama, MLX), retrieval and persistent agent memory, MCP and tool-calling, Python and TypeScript
Now
Since April 2025 I have been independent: a self-directed technical deep-dive into applied AI while caring for a new child full time. Currently building PsychBill, billing automation for Canadian psychiatrists (pre-revenue, Phase 1 in development). Open to product leadership roles where AI capability is the product: B2B SaaS adding real AI to real workflows, founding product roles, or applied-AI consultancies.