I'm Director of GTM Engineering at Tam To Target, a B2B GTM agency. I design and run the internal platform a team of six GTM engineers builds client systems on: outbound, research, signals and content for 34 client companies, mostly in EdTech, K-12, higher education and gov-tech (September 2026).
The engineers are the platform's users. It gives them:
- a catalog: one client slug resolves that client's repo, database, Slack channel, sending workspace and task board, generated from a registry, never hand-edited;
- a read-only CLI whose command set was chosen from telemetry (the 14 commands that cover about 95% of real use);
- golden paths: published skills for the common jobs (campaign copy, dial research, list upload, queue grooming);
- unattended jobs as separate apps, with drift detection between the repo and what is deployed;
- observability of every agent session and tool call;
- guardrails derived from recorded failures, and an eval-gated model allowlist;
- propose-then-apply approval for anything that writes, sends or spends.
I design the systems and run them. Coding agents write most of the code. The telemetry is how I know the result is right.
| You are | Start with | Then |
|---|---|---|
| A platform, AI-infra or devtools engineer | runtune and its evidence doc | model-eval-gate, agent-tenancy, internal-gtm-platform |
| Hiring for GTM platform, GTM systems or GTM engineering leadership | internal-gtm-platform: the layers, what they measured, the known gaps | "What it measured" below, then gtm-engineering-operating-model |
| A GTM engineer or GTM engineering leader | gtm-pipeline, gtm-research, gtm-deliverability | cc-logger |
| Layer | What it gives GTM engineers | Public reference implementation |
|---|---|---|
| Catalog and tenancy | One slug resolves every system a client has; the model never picks a tenant | agent-tenancy |
| Observability | Every agent session and tool call in one warehouse | cc-logger · codex-logger · cursor-logger |
| Policy and guardrails | Rules at the tool boundary, written from recorded failures | callusguard (components: agent-guard, codex-guard, wroteonly) |
| Harness feedback loop | Telemetry → proposed change → human approval → measured against a control | runtune |
| Model gateway | A cheaper model takes a task only after an eval clears it | model-eval-gate |
| Context and knowledge | Governed, versioned writes to shared memory | knowledge-graph-governance |
| Self-service, delivery, human approval, team interfaces | CLI, published skills, fleet drift detection, propose-then-apply loops, an authority ledger, a Slack agent | Private (described in internal-gtm-platform) |
| Workloads on the platform | List building, research, deliverability | gtm-pipeline · gtm-research · gtm-deliverability |
- I expected the platform to double how many clients each engineer could carry. It did not. From Q1 to Q3 2026, clients with a launched campaign per engineer rose from 5.4 to 6.8 while launches per engineer stayed at about 14 a month. Headcount and process changed in the same months, so these numbers do not show the platform caused the rise. A baseline for the next quarter is being taken now.
- Emails sent rose 1.81x over the same period (58k to 105k a quarter).
- Cold-email rules are graded against 171,151 sends and 3,002 human replies. About 74% of raw "replies" were autoresponders, so dashboard reply rates run roughly 4x high.
- Dial research: 100 organisations in 14.2 minutes; 0 fabricated contacts in a 1,472-record evaluation.
- runtune was developed against about 238,000 recorded tool calls and model requests. One of its guard rules made failures worse until it was rewritten, and five of its recommendations changed once checked against production. The evidence doc keeps both.
Client data, credentials, provider adapters and company-specific policy. The public repos are reference implementations extracted from the running platform.



