Sources and provenance

Every source this course draws on, ranked by tier, with the disqualification reasoning applied to each named practitioner and vendor

6 min read

This course was researched fresh via web search in August 2026 — there was no existing brief to compile from. It uses a six-tier source ranking (primary research, down through researcher, researcher-practitioner, disclosed-numbers operator, press/company, and surface content) and runs a disqualification check on every named practitioner or vendor before citing them: does the source engage with evidence rather than cite it decoratively, does their confidence stay in proportion to what's actually known, are they speaking inside their real expertise, and — for anyone claiming operator status — do they actually disclose numbers rather than just implying a track record.

Tier 1 — Primary research

Peng, Kalliamvakou, Cihon, Demirer (Microsoft Research / MIT), "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot" (2023, arXiv 2302.06590) — a genuine randomized controlled trial with disclosed methodology, sample, and task design. Used in AI-assisted development: a reality check for the 55.8%-faster finding on a narrow, well-specified coding task.

METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (2025, arXiv 2507.09089) — a genuine RCT, 16 experienced developers, 246 real tasks, disclosed methodology including the authors' own caveats about residual experimental artifacts. Used in the same lesson for the 19%-slower finding on realistic, existing-codebase work, and for the perception-versus-reality gap.

US Bureau of Labor Statistics, Business Employment Dynamics program — establishment survival data drawn from unemployment-insurance tax records covering the full private sector, not a survey. Used in Survival rates: what the data actually shows for the 34.7%-at-ten-years, ~78%-at-one-year, ~51%-at-five-years figures, and for the Information-sector historical underperformance pattern (flagged in that lesson as dated for the sector-specific figure, current for the all-industry figures).

Tier 2–3 — Researcher / researcher-practitioner

David Skok (partner, Matrix Partners; author, "SaaS Metrics 2.0," ForEntrepreneurs.com) — Tier 3. Passes the disqualification check: he derived the LTV:CAC 3:1 and CAC-payback benchmarks from his own venture portfolio's real numbers rather than citing a study decoratively, and — critically — he explicitly scoped the heuristics to mature, venture-backed B2B companies at steady state and warned against applying them indiscriminately, which is the opposite of selling false certainty. Used in The unit economics math. His blind spot, named directly in that lesson: the benchmarks were calibrated to his own portfolio circa 2010–2011, not to early-stage or non-venture-backed companies, and most of the content repeating "3:1" today has stripped that scoping out — which is the repeaters' failure, not Skok's.

Shikhar Ghosh (senior lecturer, Harvard Business School) — Tier 2–3. A real institutional researcher with a substantial disclosed sample (2,000+ venture-backed companies, 2004–2010). Held at Directional rather than Established in Bootstrapping vs fundraising for two reasons named there: the widely-cited 75% figure reached the public through a 2012 Wall Street Journal interview rather than a published paper available for independent scrutiny, and the sample is now dated relative to the current venture market.

Tier 4 — Operator (disclosed numbers)

Rob Walling (co-founder, MicroConf and TinySeed; founder, Drip.com — acquired by Leadpages, publicly disclosed) — Tier 4. Passes the check on real disclosed track record (a real exit, a real ongoing accelerator portfolio) rather than an implied one. Used in The distribution problem for the "biggest risk is that no one cares" framing — held as practitioner judgment from pattern-matching across many companies, not as a measured finding, because it isn't one and Walling doesn't present it as one.

Baremetrics Open Benchmarks — real aggregated billing data from 800+ companies using the product, including a disclosed 410%-median-ROI figure for its payment-recovery feature from a specific, dated (December 2024) cohort of 148 customers. Used in Real, disclosed benchmarks. Held at Directional rather than Established for thin disclosure of the underlying sample's composition, and a direct commercial interest in the benchmark tool looking authoritative.

SaaS Capital — a specialty lender to SaaS companies, running a disclosed 1,000+-company annual survey for fifteen consecutive years. Used across Real, disclosed benchmarks and Bootstrapping vs fundraising for growth-rate, NRR/GRR, and spending-by-function figures. Held at Directional: real sample size disclosure, but a lender with a direct interest in the SaaS-as-investable-asset-class narrative the report reinforces.

ChartMogul — subscription-analytics company, aggregated billing data cross-referenced with Dealroom funding classifications. Used in the same two lessons for the bootstrapped-vs-VC growth comparison. Held at Directional: real aggregated data, but no disclosed sample size and a direct interest in the analytics product the data is a byproduct of.

Tier 5 — Press / company (commercial interest, self-reported or promotional)

High Alpha / OpenView SaaS Benchmarks survey — 800+ self-reported respondents, run by venture-capital firms. Used in Real, disclosed benchmarks for the ~110% NRR figure, held at this tier for self-reported (not independently verified) data plus the running firms' interest in the category.

Bessemer Venture Partners — origin-story amplifier and, in 2024, reviser of its own "Rule of 40" into a "Rule of X." Used in The recurring revenue mechanism, explicitly flagged there as a firm revising the heuristic it popularized, which is worth reading with that self-interest in mind rather than as independent validation.

Chargebee — billing-infrastructure vendor, "State of Recurring Revenue and Monetization" 2025 survey. Used in Pricing models for the hybrid-pricing-adoption trend, flagged directly for selling the infrastructure that makes hybrid pricing easier to implement.

Gartner — low-code/no-code market-adoption forecast, reaching this research only through press coverage of a paywalled report. Used in No-code vs custom-built, flagged as a forecast about enterprise application development broadly, not a measurement specific to solo-founder SaaS products.

Startup Genome, 2011 report — a real, published study of 3,200+ high-growth technology startups, named specifically in Survival rates: what the data actually shows as one likely origin of the "90% fail" figure that circulates — and immediately disqualified as a source for that specific claim, because its self-selected high-growth-cohort sample cannot support a general failure-rate statistic, which is exactly the kind of context-stripping this course's methodology exists to catch.

Tier 6 — Surface content, named and discarded

Every specific percentage version of "X% of SaaS startups fail" (90%, 92%, 95%, and close variants) circulating across SaaS content — named and explicitly discarded in Survival rates: what the data actually shows as untraceable to any disclosed, current methodology, regardless of which authority they're attributed to.

Specific dollar-figure claims for "cost to build a SaaS MVP" (no-code versus custom-development cost comparisons citing precise ranges) — named and discarded in No-code vs custom-built as unsourced content-marketing estimates from parties selling one option or the other.

Specific "time to $10K MRR" milestone timelines and per-founder revenue figures (including a widely-repeated "$500/month median micro-SaaS" claim) — not used anywhere in this course's lessons; named here only so you discount them if you encounter them elsewhere, on the same "repeated until it sounds authoritative, never actually sourced" pattern as the failure-rate figures above.

What this means for how to use this course

Every [Directional] tag in this course traces to a real, disclosed source with a stated interest, listed above — not to an unnamed "studies show." Every [Established] tag traces to either a Tier 1 primary source or an algebraic identity true by its own definition. Nothing in this course rests on a Tier 5 or 6 source without that source's specific limitation named at the point of use. Where this research hit a genuine gap — most visibly, the absence of any disclosed, methodologically-transparent SaaS-specific survival-rate figure at any tier — that gap is stated as a gap, not filled with the most convenient number available.

SaaS · progress saved in this browser · sign in to sync across devices

That’s the end of SaaS.

You've finished the reading order. 13 lessons left unmarked — worth a pass before you call it done.

Back to the contents