an agent skill · a gate before the playbook — onboarding, support, health scores, renewals

/success

success keeps the customers you already have. Ask it to plan onboarding and it writes the adoption program, not just an in-app checklist; ask it whether your health score is trustworthy and it checks the math before it trusts the color; ask whether you even need a customer-success function yet and it answers honestly — often "not yet," with a named pointer to whatever already covers the gap. Onboarding, help content, support process, AI-support rules, renewals, and the customer's side of an incident — 15 references and four self-testing calculators behind one router. What makes it different: most churn-and-health advice repeats a number nobody ever actually measured. This pack names which of those numbers have never been published anywhere — and builds your plan without needing one.

# natural language — no flags, no fixed pipeline /success our beta customer's health score just flipped from green to red — is that signal or noise?

runs onClaude CodeCodexCursorAntigravityopencodeGrok BuildHermes

The gate, then the router. There is no fixed onboarding-to-renewal pipeline to run start-to-finish — each job stands alone and enters where your request is, after the gate clears. The animation traces one path (the gate, then the flagship's honesty check on a health score); the sections below map the whole surface it routes across.

Hold the relationship, honestly

Keep the customer relationship after the sale, and be honest about what you can actually know about it. success owns the post-purchase jobs — the onboarding program, help content, the support process, AI-support policy, health and risk signals, voice-of-customer interpretation, retention execution, renewals and expansion, and the customer's side of an incident — and it owns the epistemics that come with them: which of those signals can carry a decision, and which cannot. success writes programs, policies, taxonomies, and workflows. It writes no production code, no experiments, and no numbers it cannot source.

success owns

everything between "does this customer stay" and "what can we actually prove about that"

  • the existence gate — whether a distinct success function is worth building yet, before a single health score or playbook exists
  • the programs — onboarding as an adoption plan, help content and the knowledge base, support process, SLAs as promises made to a real person
  • health and risk signals ⭐ — whether a score or model can actually predict anything, and whether the churn event is even observable in the first place
  • the relationship motions — AI-support policy, voice-of-customer interpretation, retention execution, renewals and expansion, and the customer's side of an incident

hands off to

the "Not this skill" table — twelve asks this skill declines by design

  • design — the in-product onboarding flow, NPS's definition and survey method; success owns the onboarding program and the score's interpretation, not the widget or the question
  • growth — whether a retention or win-back mechanic works; success runs the motion a growth test already validated, as an ongoing relationship, not a live experiment
  • product — pricing tiers, the commercial-lifecycle model, dunning and billing-failure recovery; success operates what product models
  • ai — the support assistant's turn loop, retrieval, and evals; success sets the escalation condition and the resolution definition, ai measures compliance with them
  • operate — internal SLOs, on-call, incident execution; success owns the customer-facing response-time promise, never the internal target
  • marketing · automation · sales · data · quality · frontend/backend/ai — acquisition and the store listing, the approval gate on an irreversible action, negotiating an expansion deal (provisional), the tracking plan and PII governance, build-defect triage, and any production code at all

The default reader is a builder with users, not a CS team with a book of accounts. So the front door is a gate, not a playbook — because no source in the CS canon ever says when a distinct success function becomes worth having. Where the enterprise canon still teaches something useful, this pack lifts its epistemics — input-expiry rules, segmented thresholds, treating missing data as risk, tie-breakers on a category list — and declines its machinery — CSM ratios, QBR cadence — with reasons stated, never by omission.

The existence gate — the front door

Before a single onboarding plan, health score, or escalation policy gets written, this pack asks whether a distinct success function is worth building yet. No source in the CS literature — Gainsight's own book, the CSM-ratio literature, any vendor playbook — states when that becomes true; the whole canon assumes the function already exists and only asks how to staff it. The Existence Gate uses the one practitioner sentence that makes it checkable, from Arvid Kahl, a solo founder: "Customer Success (compared to mere customer support) is so much more important when your revenue scales with your customer's revenue."

testverdictwhat it means
Customer growth changes what they pay youappliesSeat expansion, usage-based pricing, revenue share — watching account health and driving adoption pays for itself, because a bigger customer is a better customer.
Flat-rate self-serve — success doesn't change the billlighter-weightOptional until it visibly earns its cost. growth, marketing, and operate can cover most post-purchase needs in the meantime.
Unsettled, or the request assumes CS exists without justifying itnot yetNever invented silently — the verdict names which sibling already covers the gap: marketing for retention communication, growth for retention experiments, operate/product for incident comms and churn as a guardrail metric.
Contested, shipped as contested: the research found the solo-scale replacement for CS process is pricing first, then AI tier-1 support — not documentation, which this pack's own original hypothesis assumed and two research passes failed to support. Both practitioners cited land on pricing as the lever; no opposing voice was found arguing the reverse. That one-sidedness is presented as a strong position with the absence noted, never as a settled fact — and it never excuses a support residue automation can't close: a customer who genuinely misremembers what they bought is a memory problem, not a lookup problem, and the gate's verdict says "lighter," never "zero."

The ten primary jobs

Each job is one reference, read fully only when its route is selected. Health and risk signals is the flagship because it is the file every red score, every renewal risk, and every churn-save motion eventually has to answer to. This is the whole post-purchase surface, not a headline slice.

I need to…ReadContribution
Decide whether a distinct success function is worth building yet when-success-applies.md The Existence Gate. The attributed test, what covers post-purchase meanwhile, and the contested pricing-first correction over docs-as-deflection
Build an implementation/adoption plan or an education cadence outside the product onboarding-and-adoption.md The onboarding program, with the design seam stated up front; time-to-first-value consumed from the PRD, never re-derived; Murphy's current Goal + Appropriate Experience formula, dated
Decide what belongs in a knowledge base, or answer a deflection claim education-and-knowledge-base.md The family's first definition of good help content; an article-lifecycle rule; the zero-result/zero-click backlog as the observable stand-in for an unmeasurable deflection rate
Set up or audit a ticket taxonomy, SLA metric, or response-time promise support-process-design.md Two taxonomy axes plus the three converged status buckets; first/next/resolution SLA naming; the response-time promise, and the rule against any SLA default not derived from your own capacity
Define when a bot hands off, or how to read a vendor's automation claim ai-support-policy.md Escalation is symbolic, never a confidence float — nine real triggers, none numeric; resolution-rate honesty (confirmed vs. assumed); vendor claims read against the one measured instrument that exists
Build, buy, interpret, or audit a health score or churn model health-and-risk-signals.md Flagship. Score-vs-prediction as different artifacts; "not published, never not measured"; the Observability Question; four-study convergence; the honest ML ceiling; the reframe — a score plus a mandatory gap and action
Read an NPS/CSAT/CES score or a feedback trend honestly voice-of-customer.md Interpretation only — three siblings already define the metric; the Display-vs-Response denominator check, the NPS-growth and CES-superiority replication failures
Run a churn-save, win-back, or renewal-nudge motion a test already validated retention-execution.md The reserved job name and the third box: standing is operate's, experiment-bound is growth's, a validated motion held as a relationship is success's; target by uplift, not raw risk rank
Run a renewal, evaluate an expansion signal, or report an NRR/GRR figure renewals-and-expansion.md Four audited NRR definitions that are never compared across each other; expansion signals labeled by evidence grade; advocacy stops at qualification and consent
Communicate with customers during an incident or a decommission incident-and-account-comms.md Co-owned with product, always — operate states fact, impact, and timing; success manages the relationship and originates no technical claim

Full router table & invariants: SKILL.md.

Three surfaces + one additive overlay

At most one base surface reshapes every job for how the product is actually run — the same onboarding job resolves differently for a founder answering their own email than for a CSM-staffed account. The agentic overlay is additive — it stacks on top of whichever base you picked, never replaces it — and carries the violet identity throughout this page, the same convention automation, operate, quality, data, marketing, and growth use for their own additive overlays.

Self-serve solo surface-selfserve-solo.md

The default: no dedicated CS function exists — one founder or a small team is the entire success operation. If a request names no business model, this is what the skill assumes — and it says so rather than stalling for a clarification it doesn't need.

reshapesautomation and AI tier-1 support as the team, not a documentation library · which jobs are done by hand versus automated

B2B high-touch surface-b2b-high-touch.md

A named CSM, a service tier, or an account-level health score exists. The canon's home turf — this surface lifts its epistemics (input expiry, segmented thresholds, ignorance-as-a-finding) and declines its machinery (CSM ratios, QBR cadence), with reasons stated.

reshapesthe health-score audit against a named public exhibit's five failure modes · the renewal conversation's sales seam

Mobile subscription surface-mobile-subscription.md

A native iOS/Android subscription — store reviews are a support channel, a cancellation path the developer doesn't control, and billing retry the platform runs itself.

reshapesapp-store-review responses into the existing taxonomy · the pre-cancel in-app prompt within store constraints

Agentic additive

A model decides and acts on support, health, or retention without a human reviewing that instance — not a human approving one draft a model produced. Stacks on, does not replace. Escalation triggers and the resolution definition don't loosen because a model runs the loop; they get more load-bearing, not less.

reshapeswhich decisions are model-decided vs. human-approved per instance · what's ai's (turn-loop mechanics) vs. success's (the policy envelope around them)

Health and risk signals — the flagship

Ask an agent to build a churn model and it either ships a color with no math behind it, or a number with no source. This pack checks first: can a score or model actually predict anything, and is the churn event even observable in the first place — before either ships. Even the most transparent public customer-success program in the research, GitLab's own handbook, does not ask its health score to predict churn: the score is one artifact, prediction is two separate internal models, and neither publishes an accuracy figure.

"Not published, never not measured" — a headline number, checked

What nobody discloses, at three levels

  • score ≠ predictionGitLab's own health score doesn't predict churn — that's two separate, unpublished-accuracy internal models (PTE/PTC), even at the most instrumented public exhibit
  • nobody disclosesfive named CS vendors, the entire open-source repo corpus, and GitLab's own in-house data-science team — zero of them publish a churn-model accuracy figure against real outcomes
  • the Observability Questionin a contractual setting (subscription, renewal date) churn is observed on a known day; in a non-contractual one (usage-based, self-serve) it's unobserved — the CS canon never draws this line

Four independent academic studies converge on the same ceiling: risk-ranking isn't response-ranking (Ascarza 2018 — the highest-risk decile improved least); tenure alone (r=.199) out-predicts every survey metric, at a ceiling of about 4% of variance (de Haan 2015); feedback metrics predict the present, not the future (van Doorn 2013) — see health-and-risk-signals.md §4 for the full table.

A headline number, checked

93–94% accuracy — the number that circulates. Traced to resampling the full dataset before the train/test split, or a majority-class baseline dressed as a result 85.5% baseline — what "nobody churns" already scores on the same public dataset, no model required 78–81% accuracy · AUC 0.83–0.85 — the clean, independently reproduced ceiling across three non-leaky implementations. This is what a churn model honestly achieves
Score-vs-prediction are different artifacts. A health score alone next to a renewal date is folklore; a score paired with a stated gap — which input is stale, which measure is off, why it's red — and a required action is a prioritization device. Josh Pigford, founder of a company whose product was churn analytics, publicly could not attribute his own 2× churn improvement — against every incentive to claim credit for it.
a bare percentage with no method is unusable — including this pack's own best-graded anchora red score is a rubric override, not an arithmetic average, for a customer's first 30 days

What transfers from the enterprise canon, at no cost: input expiry (show nothing rather than stale data), segmented thresholds (a measure switched off for a tier, not scored badly), treating missing data as risk rather than neutral, and a tie-breaker rule for every category boundary — none of it requires the CSM ratios or QBR cadence it usually ships wrapped in.

What makes this different

Customer-success advice is not scarce — it's fragmented. Real tools cover real pieces of this surface; none of them integrate, and almost none of them cite a source for the numbers they repeat. What doesn't exist anywhere else is a pack that puts the whole post-purchase relationship and the honesty layer underneath it in the same repo, and grades every figure it leans on instead of repeating it because it sounds measured.

Real tools, zero overlap, nobody integrates

One popular skill covers churn prevention with real reach but only two of the twelve canonical CS jobs, citing nothing. One official pack is well-built but purely tactical — triage, escalate, one KB article — with no onboarding, health, or renewal surface at all. The most architecturally complete pack found cites zero external sources across nearly 29,000 lines, and is built for enterprise CSM teams, not a builder with users.

  • the gate is asked first — whether success applies at all, before the apparatus, not discovered as an afterthought once a health score already exists
  • the honesty layer travels with the score — what a health score can predict, why predicting correctly may still not help, and whether the churn event is observable at all, answered with evidence rather than left as an open question
  • seams stated, not assumed — onboarding program vs. in-product flow, retention execution vs. experimentation, policy vs. mechanics — each boundary is named and cited, never silently claimed

Every figure carries its evidence grade, or it doesn't ship

Customer-success canon repeats a small set of numbers no primary source actually supports. This pack traced each one back and built a rule against shipping it again.

  • "5% retention lifts profit 25–95%" — appears in no primary source; never shipped, the chain is taught instead
  • NPS as "the best predictor of growth" — failed replication (a peer study's own metric beat it 2 of 3 times); cited with the replication, never repeated alone
  • "96% of unhappy customers never complain" — a worst-case slice of one study sold as a universal law; the real range runs roughly 5–75% depending on what's at stake
  • calculators are executable, not prose — nrr_calc.py, churn_baseline_check.py, and sla_capacity_calc.py each self-test against a published anchor, never a hand-built formula this page would then repeat after it drifted
Fragmented, not empty. Real prior art exists — a churn-prevention skill with genuine reach, an official support-triage pack, an architecturally deep CSM library — and is described respectfully throughout. The gap this pack fills is that no one of them covers onboarding through renewal and grades its own numbers. An honest "this has never been measured, and here's why" is a deliverable, not a failure to produce one.

The universal invariants

These govern every route, whichever references it loads.

What a pass produces

Gate verdicts, programs, taxonomies, policies, score audits, and workflows are delivered as documents. None is delivered as production code. A success artifact is incomplete unless it carries the gate verdict (or a statement that success already applies), every figure with its source and evidence grade, the seam naming which sibling owns the adjacent half, and what the artifact cannot support.

step 1run the gatewhether success applies at all; then the one primary job and the base surface, overlay only if a model acts unreviewed
step 2name the decisionwhat's actually being decided, and who owns that decision once the artifact exists
step 3consume upstreamthe PRD's TTFV, design's onboarding flow and NPS definition, growth's readout, operate's stated fact — cited, never re-derived
step 4grade every figureevery benchmark or "best practice" relied on, and where its evidence stops
step 5produce the artifacta gate verdict, program, taxonomy, policy, score audit, motion, or workflow — with a named proof source behind every claim
step 6run the honesty checksthe Observability Question for anything predictive, capacity for anything promised, voluntary/involuntary for anything about churn
step 7hand offstate what the artifact cannot support, and emit a compact handoff when downstream work is expected

nrr_calc.py

assets/nrr_calc.py

The four audited NRR definitions as selectable modes, never blended into one number — the SEC-verbatim wording, self-tested against a reference implementation.

churn_baseline_check.py

assets/churn_baseline_check.py

Checks a headline accuracy claim against the naive majority-class baseline, plus a concordance sanity band that flags the same leakage this page's flagship names.

sla_capacity_calc.py

assets/sla_capacity_calc.py

What response time your actual coverage can honestly promise, and whether ticket volume already exceeds it — no copied template number.

health_score_audit.md

assets/health_score_audit.md

The ten-row checklist a score must pass before anyone trusts a color — the only non-runnable asset, deliberately, since a score's own inputs need a human read.

Eval suites

evals/activation · traversal · output · compression-ablation

Does the router reach the smallest sufficient reference set; does the content change model behavior on real gate/health-score prompts; does a never-ship figure ever resurface, even inside a correction of it.

The handoff seams

success operates independently when invoked alone, and uses compatible upstream artifacts without silently overriding them. handoff.md maps every seam between this pack and the rest of the family from success's side — including three rows no sibling has drawn yet, because success ships after product, operate, and automation did.

success emits

gate_verdict: <applies /
  lighter-weight / not-yet>
job_and_surface: <which one,
  and any surface assumed>
claims: [claim, named proof source,
  evidence tier]
seam: <sibling pack + filename
  owning the adjacent half>
cannot_support: <what the
  artifact does not prove>

marketing

Retention and churn signal that feeds messaging and lifecycle work. Success consumes acquisition-side content for onboarding continuity — marketing's lifecycle job stops at acquisition; onboarding education and retention communication are success's, stated flatly, as marketing itself states it.

growth (provisional)

A validated retention motion, sent — the third box success draws from its own side: standing is operate's, experiment-bound is growth's, a validated motion held as an ongoing relationship is success's. Success consumes growth's experiment readout and never reports it as its own finding.

ai

Policy artifacts — guidelines, exclusion edges, escalation policy, deflection-rate honesty. Success consumes the support-assistant surface mechanics and its evaluation; ai owns mechanics, success owns policy, cited by pack and filename on every use.

Three rows this pack originates. product, operate, and automation each already cede work to success in their own prose — but none carries a reciprocal success row, because success shipped after they did. This file states those seams from success's side: product hands success TTFV and pricing terms and receives support/satisfaction/churn evidence back; operate hands success the fact, impact, and timing of an incident, and success turns it into customer-facing comms, co-owned with product, never alone; automation's refund-approval flagship is a support workflow end to end — success owns the process and the refund policy, automation owns the approval gate.

The rest of the family — each an independently installable pack with its own guide:

Start here

Install once. It's a plain SKILL.md router — no flags, no config, no fixed pipeline — so it activates on natural-language phrasing ("do we need customer success yet," "audit our health score before we trust it," "draft our AI-support escalation policy") rather than a fixed command.

# skills.sh ecosystem npx skills add gabros20/success-skill -g -y # clone + manual copy git clone https://github.com/gabros20/success-skill cp -R success-skill/skills/success ~/.claude/skills/success # use — natural language, any host /success do we need a distinct customer-success function yet, or is growth/marketing covering it fine /success audit our health score before we trust the red/yellow/green colors /success draft our AI-support escalation policy — when does the bot hand off to a human

The same install runs on any Agent Skills host. Codex triggers with $success; manual copy into any client's skills directory also works.

what's in the repo
skills/success/ the skill: SKILL.md (router) + 15 references/ + 4 self-testing assets/ + evals/ research/ multi-channel research corpora + build-gate synthesis site/ this guide — deploys to successskill.vercel.app README.md · SOURCES.md · LICENSE

More detail: SKILL.md · SOURCES.md — source attribution, licensing rule, and the numbers this pack refuses to ship.