7+ years in AI product strategy, ML and AI engineering, hands-on
○ AWAITING REVIEW
20 yrs shipping · Principal, agentic AI strategy
I own agentic AI strategy at Character.ai and set the org-wide standards for AI tooling — while still writing the systems myself.
Twenty years underneath it: Coinbase, Netflix, Google Cloud, VMware.
▸ TAP FOR EVIDENCE · +4% SYNC
3+ years of experience in the game development industry
○ AWAITING REVIEW
4 shipped products, real players. No studio badge.
My route was not a game studio, and I would rather say so than stretch a date range.
Case Files — timed detective game, Character.ai Interactive Originals. StoryLab — LitRPG audio. c.ai/studio — weekly microdrama. AI She Wrote — built and operated alone.
If the requirement is studio tenure, I do not meet it. If it is shipping games that agents run, I do.
▸ TAP FOR EVIDENCE · +4% SYNC
Product management: user needs, roadmaps, production workstreams
○ AWAITING REVIEW
Owns the roadmap. From the builder's seat, not a PM's.
I define what gets built for agentic AI at Character.ai and run the workstreams that ship it, which is the 'Builder PM' half of this role. What I have not done is carry a formal PM title, run structured user research, or own a roadmap negotiated across a product org. I have prioritized by shipping and reading telemetry, not by writing PRDs.
▸ TAP FOR EVIDENCE · +4% SYNC
Games or interactive entertainment — shipped features, game engines (Unreal, Unity), or AI experiences for players
○ AWAITING REVIEW
AI for players: 4×. Unreal/Unity: no.
This requirement is an or, and I clear the third branch decisively — building AI experiences for players is the whole archive above. I have not worked in Unreal or Unity, and I am not going to imply otherwise. My interactive surfaces are web and app clients over agent backends. If the work is engine-side gameplay programming, that is a ramp for me; if it is the agent layer that talks to the engine, that is what I already build.
▸ TAP FOR EVIDENCE · +4% SYNC
Deep, practical agentic AI in production — multi-step reasoning, tool use, multi-agent orchestration
○ AWAITING REVIEW
AgentX — 30+ agentic turns, 20M+ MAU, no human in the loop
A planner agent decomposes the request, executor agents generate image, video, and audio, and reviewer agents gate every stage on trained eval scores. 30+ agentic turns, 6 specialized tools, no human in the loop after the user submits a prompt.
▸ TAP FOR EVIDENCE · +4% SYNC
Strong Python and production experience with agentic frameworks (LangChain, LangGraph, AutoGen, Google ADK)
○ AWAITING REVIEW
Python for post-training and evals. Go for the runtime, by choice.
Evals, judge training and the annotation loop are Python.
The runtime is not. I open-sourced claude-agent-sdk-go instead of adopting LangChain — at 20M+ MAU I wanted bounded goroutine pools, real context propagation, a static binary.
I know those frameworks and chose against them. A house language is a preference, not a skill gap.
▸ TAP FOR EVIDENCE · +4% SYNC
Designing and operating evaluation infrastructure — offline benchmarks, automated scoring, online experimentation
○ AWAITING REVIEW
JudgeJudy + Judge Dredd — offline harness, online loop
Open-sourced a multimodal eval harness (AI judges, trained scoring models, human scoring pipeline), then built the self-improving layer: Judge Dredd pulls Amplitude and Statsig telemetry over MCP, correlates eval scores against real user behavior, and feeds validated gates back into the runtime. Every AgentX release clears the wall before players see it. This is the requirement I would point at first.
▸ TAP FOR EVIDENCE · +4% SYNC
Deep GenAI ecosystem understanding — open models, infrastructure, safety, hosted vs in-house trade-offs
○ AWAITING REVIEW
Post-trained (SFT + RL) open models in-house. Frontier models in the pipeline.
I post-train open models — SFT and RL — for specific agentic use cases, so the experiences serve at scale and quality rather than paying frontier prices for every call.
The judge runs on an 8B open-weight VLM because it lives at the release gate and has to keep up with CI. We A/B'd a larger model: marginally better, significantly slower.
The generation path uses frontier models. That hosted-versus-in-house call, made on latency and cost rather than benchmark vanity, is the trade-off this bullet is asking about.
▸ TAP FOR EVIDENCE · +4% SYNC
Research prototypes to production systems, cross-functional and fast
○ AWAITING REVIEW
Every system here started unassigned. I own the pager after.
I scope it, build it, ship it, and own the pager afterwards. AgentX, the eval wall, the SDK, and Clauder all began as prototypes and are now load-bearing. I work best where the constraint is the problem itself rather than the process around it.
▸ TAP FOR EVIDENCE · +4% SYNC
Strong communicator across technical and non-technical stakeholders
○ AWAITING REVIEW
Conference talk, Confluent podcast, org-wide standards
The talk and the podcast are both in TRANSMISSION above — unedited, judge me on them. Internally I set the org-wide standards for AI tooling at Character.ai, which is mostly a persuasion problem rather than a technical one.
▸ TAP FOR EVIDENCE · +4% SYNC
Managing a small team of engineers
○ AWAITING REVIEW
Two orgs grown to 20+ engineers at Coinbase
Two orgs to 20+ engineers, company-wide technical strategy for Datastores, and a vendor replaced in-house at 90% cost savings. A small team is a smaller version of a problem I have run before.
▸ TAP FOR EVIDENCE · +4% SYNC