Skip to content
View kkrlstrm's full-sized avatar

Block or report kkrlstrm

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
kkrlstrm/README.md

Kai Karlstrom

Platform engineering for GTM engineers.

I'm Director of GTM Engineering at Tam To Target, a B2B GTM agency. I design and run the internal platform a team of six GTM engineers builds client systems on: outbound, research, signals and content for 34 client companies, mostly in EdTech, K-12, higher education and gov-tech (September 2026).

The engineers are the platform's users. It gives them:

  • a catalog: one client slug resolves that client's repo, database, Slack channel, sending workspace and task board, generated from a registry, never hand-edited;
  • a read-only CLI whose command set was chosen from telemetry (the 14 commands that cover about 95% of real use);
  • golden paths: published skills for the common jobs (campaign copy, dial research, list upload, queue grooming);
  • unattended jobs as separate apps, with drift detection between the repo and what is deployed;
  • observability of every agent session and tool call;
  • guardrails derived from recorded failures, and an eval-gated model allowlist;
  • propose-then-apply approval for anything that writes, sends or spends.

I design the systems and run them. Coding agents write most of the code. The telemetry is how I know the result is right.


Start here

You are Start with Then
A platform, AI-infra or devtools engineer runtune and its evidence doc model-eval-gate, agent-tenancy, internal-gtm-platform
Hiring for GTM platform, GTM systems or GTM engineering leadership internal-gtm-platform: the layers, what they measured, the known gaps "What it measured" below, then gtm-engineering-operating-model
A GTM engineer or GTM engineering leader gtm-pipeline, gtm-research, gtm-deliverability cc-logger

The platform, by layer

Layer What it gives GTM engineers Public reference implementation
Catalog and tenancy One slug resolves every system a client has; the model never picks a tenant agent-tenancy
Observability Every agent session and tool call in one warehouse cc-logger · codex-logger · cursor-logger
Policy and guardrails Rules at the tool boundary, written from recorded failures callusguard (components: agent-guard, codex-guard, wroteonly)
Harness feedback loop Telemetry → proposed change → human approval → measured against a control runtune
Model gateway A cheaper model takes a task only after an eval clears it model-eval-gate
Context and knowledge Governed, versioned writes to shared memory knowledge-graph-governance
Self-service, delivery, human approval, team interfaces CLI, published skills, fleet drift detection, propose-then-apply loops, an authority ledger, a Slack agent Private (described in internal-gtm-platform)
Workloads on the platform List building, research, deliverability gtm-pipeline · gtm-research · gtm-deliverability

What it measured

  • I expected the platform to double how many clients each engineer could carry. It did not. From Q1 to Q3 2026, clients with a launched campaign per engineer rose from 5.4 to 6.8 while launches per engineer stayed at about 14 a month. Headcount and process changed in the same months, so these numbers do not show the platform caused the rise. A baseline for the next quarter is being taken now.
  • Emails sent rose 1.81x over the same period (58k to 105k a quarter).
  • Cold-email rules are graded against 171,151 sends and 3,002 human replies. About 74% of raw "replies" were autoresponders, so dashboard reply rates run roughly 4x high.
  • Dial research: 100 organisations in 14.2 minutes; 0 fabricated contacts in a 1,472-record evaluation.
  • runtune was developed against about 238,000 recorded tool calls and model requests. One of its guard rules made failures worse until it was rewritten, and five of its recommendations changed once checked against production. The evidence doc keeps both.

What stays private

Client data, credentials, provider adapters and company-specific policy. The public repos are reference implementations extracted from the running platform.

Pinned Loading

  1. gtm-pipeline gtm-pipeline Public

    Reference implementation of a list-building path for GTM engineers: a plain-English campaign brief becomes a deduped, qualified, enriched, sequencer-ready contact list in seven stages, with swappab…

    Python 16 1

  2. cc-logger cc-logger Public

    Observability for Claude Code: HTTP hooks record sessions, sub-agents and tool calls to Postgres, so agent work can be inspected, compared across runs, and turned into guardrails.

    Python 11 1

  3. runtune runtune Public

    A runtime learning loop for coding agents. Turns Claude Code, Codex, Cursor, Antigravity and model-router runs into human-approved, measured changes to the agent harness: guard rules, skills, sub-a…

    Python 9

  4. internal-gtm-platform internal-gtm-platform Public

    Reference architecture of an internal GTM platform: the platform one person runs for a team of GTM engineers delivering to 34 clients. Ten capability layers, the invariant each enforces, what it me…

    2

  5. model-eval-gate model-eval-gate Public

    Model gateway policy for agents: a cheaper model may take a task only after an eval on real task data clears that task, and regression re-checks catch drift. Extracted from an internal GTM platform…

    Python 2

  6. agent-tenancy agent-tenancy Public

    Tenant catalog and scoped capabilities for multi-tenant agent runtimes: a generated registry resolves each tenant's repo, database and channel before the agent runs, so the model never chooses a te…

    Python 1