High Signal Podcasts Evidence ledger
Method
Browse

Methodology

High Signal Podcasts indexes public statements as claims. A published row always has a verbatim excerpt that exists in the stored transcript segment, a date, a source, and an explicit attribution status. When the transcript does not establish the speaker, the row says so and is excluded from person counts. A timed YouTube link appears only when that excerpt came from captions on the same video; publisher and RSS transcript clocks are not assumed to match a separate YouTube edit.

The pipeline

Eight stages, each one a gate.

Every published claim passes through all eight stages in order. A claim that fails a gate does not get published at a lower confidence anyway — it stops at that stage.

  1. 01
    Discover

    Daily GitHub Actions read podcast RSS as the primary catalog; the YouTube channel feed enriches matching episodes by title and date. An unmatched YouTube upload is admitted only when RSS returned nothing or a show is explicitly configured to use YouTube as a source.

  2. 02
    Transcripts

    A canonical speaker-labelled publisher page is used first where one is supported, then RSS podcast:transcript files, then YouTube captions, then opt-in Whisper. Publisher pages are accepted only after domain, redirect, title, and structure checks, so an operational failure stays retryable and only a checked, genuine absence is recorded as no_transcript.

  3. 03
    Segment

    Each transcript is cut into roughly 3,000-character windows at cue and speaker boundaries. A transcript replacement is exact: stale trailing segments are removed, and it is rejected outright once a published claim already references that episode.

  4. 04
    Attribute

    Publisher speaker labels become people only through manually reviewed label mappings; every other turn stays unknown. A row whose speaker the transcript does not establish says so explicitly and is excluded from person counts rather than guessed.

  5. 05
    Extract

    Cheap triage keeps recommendations, positions, predictions, evaluations, explanations, and commitments, and drops questions, filler, ads, and context-dependent fragments. Deterministic rules classify complete claims; a batched extraction model classifies the ambiguous remainder. Neither path invents or rewords the excerpt — the stored assertion is the same source quote.

  6. 06
    Validate

    Before a claim can be judged, its quote is looked up verbatim inside the stored transcript segment text. A claim without an exact match never advances past draft, so it never reaches the public index.

  7. 07
    Judge

    Extraction confidence and speaker confidence are banded into low, medium, and high. High-confidence, quote-validated claims publish immediately; medium-confidence claims are held for review; low-confidence and unknown-speaker claims stay in draft. A high-confidence quote with an unresolved speaker can publish labeled speaker_unverified, but it is excluded from person and distinct-recommender counts.

  8. 08
    Publish

    A published row keeps its speaker, verbatim excerpt, date, attribution status, and source link attached. The website is server-rendered from the API, so a newly published claim appears without a rebuild.

Missing evidence is shown as missing. We do not invent an answer.

Transcript excerpts can preserve errors made by a publisher or caption provider. The original episode remains the final authority, and every receipt keeps that source attached for review.

Update schedule

How often does the index update?

Two cadences: discovery is daily, extraction is weekly.

StageScheduleWhat runs
Discovery and transcriptsDaily, 06:00 UTCdiscover, then transcripts
Claim extractionWeekly, Sundays 07:00 UTCextract, all segments

An episode can appear in the catalogue within a day, while its claims stay unpublished until the next Sunday run.

The refusal conditions

What disqualifies a claim.

Enforced in judgeClaim in the API worker.

ConditionOutcome
Quote not found in stored segmentNever published - quote_not_verbatim
Quote shorter than 40 charactersCannot be anchored
Speaker is unknownNever published - unknown_speaker
Both confidences >= 0.85Published
Both >= 0.65, not both >= 0.85Held
Either below 0.65Stays a draft
Unverified speaker, extraction >= 0.85Published, labelled unverified

One published claim, traced

Worked example.

Verified live on 5 September 2026.

"I think the products that are doing this, they have a very sharp sense of how well their application is performing, and people don't talk about it, because this is their moat."

Follow it yourself: GET /api/claims/e7240228-46cd-41e6-bb59-0a67d8cb5dca returns the claim, its evidence row, and its references. The claim's assertion, its quote, and the evidence row's quote are all the same string - that identity is what "verbatim" means here.

The row carries transcriptKind: publisher_json - the publisher's own transcript, with speakers already named. That field is on every claim, and it tells you how much the attribution rests on.

Honest boundaries

What we do not claim.

  • We verify attribution, not accuracy. Whether they were right is not checked.
  • "Verbatim" means against the transcript, not the audio. Publisher mistakes survive.
  • The date is the episode's publication date. Not the recording date.
  • Timed links are conditional. Only where captions carry timing.
  • Absence is shown as absence. No generated answers.

Check it yourself

Where the data is.

Everything on this page is checkable through the documented read-only API. The website and public API currently require no account, API key, subscription or checkout - which describes the product today and is not a promise of a permanent pricing model.

GET https://api.podcasts.highsignal.app/api/stats
GET https://api.podcasts.highsignal.app/api/claims/{id}
GET https://api.podcasts.highsignal.app/api/search?q=agents

What a published claim looks like, with a real example

Search evidence