Methodology
High Signal Podcasts indexes public statements as claims. A published row always has a verbatim excerpt that exists in the stored transcript segment, a date, a source, and an explicit attribution status. When the transcript does not establish the speaker, the row says so and is excluded from person counts. A timed YouTube link appears only when that excerpt came from captions on the same video; publisher and RSS transcript clocks are not assumed to match a separate YouTube edit.
The pipeline
Eight stages, each one a gate.
Every published claim passes through all eight stages in order. A claim that fails a gate does not get published at a lower confidence anyway — it stops at that stage.
- 01 Discover
Daily GitHub Actions read podcast RSS as the primary catalog; the YouTube channel feed enriches matching episodes by title and date. An unmatched YouTube upload is admitted only when RSS returned nothing or a show is explicitly configured to use YouTube as a source.
- 02 Transcripts
A canonical speaker-labelled publisher page is used first where one is supported, then RSS podcast:transcript files, then YouTube captions, then opt-in Whisper. Publisher pages are accepted only after domain, redirect, title, and structure checks, so an operational failure stays retryable and only a checked, genuine absence is recorded as no_transcript.
- 03 Segment
Each transcript is cut into roughly 3,000-character windows at cue and speaker boundaries. A transcript replacement is exact: stale trailing segments are removed, and it is rejected outright once a published claim already references that episode.
- 04 Attribute
Publisher speaker labels become people only through manually reviewed label mappings; every other turn stays unknown. A row whose speaker the transcript does not establish says so explicitly and is excluded from person counts rather than guessed.
- 05 Extract
Cheap triage keeps recommendations, positions, predictions, evaluations, explanations, and commitments, and drops questions, filler, ads, and context-dependent fragments. Deterministic rules classify complete claims; a batched extraction model classifies the ambiguous remainder. Neither path invents or rewords the excerpt — the stored assertion is the same source quote.
- 06 Validate
Before a claim can be judged, its quote is looked up verbatim inside the stored transcript segment text. A claim without an exact match never advances past draft, so it never reaches the public index.
- 07 Judge
Extraction confidence and speaker confidence are banded into low, medium, and high. High-confidence, quote-validated claims publish immediately; medium-confidence claims are held for review; low-confidence and unknown-speaker claims stay in draft. A high-confidence quote with an unresolved speaker can publish labeled speaker_unverified, but it is excluded from person and distinct-recommender counts.
- 08 Publish
A published row keeps its speaker, verbatim excerpt, date, attribution status, and source link attached. The website is server-rendered from the API, so a newly published claim appears without a rebuild.
Missing evidence is shown as missing. We do not invent an answer.
Transcript excerpts can preserve errors made by a publisher or caption provider. The original episode remains the final authority, and every receipt keeps that source attached for review.
Update schedule
How often does the index update?
Two cadences: discovery is daily, extraction is weekly.
| Stage | Schedule | What runs |
|---|---|---|
| Discovery and transcripts | Daily, 06:00 UTC | discover, then transcripts |
| Claim extraction | Weekly, Sundays 07:00 UTC | extract, all segments |
An episode can appear in the catalogue within a day, while its claims stay unpublished until the next Sunday run.
The refusal conditions
What disqualifies a claim.
Enforced in judgeClaim in the API worker.
| Condition | Outcome |
|---|---|
| Quote not found in stored segment | Never published - quote_not_verbatim |
| Quote shorter than 40 characters | Cannot be anchored |
| Speaker is unknown | Never published - unknown_speaker |
| Both confidences >= 0.85 | Published |
| Both >= 0.65, not both >= 0.85 | Held |
| Either below 0.65 | Stays a draft |
| Unverified speaker, extraction >= 0.85 | Published, labelled unverified |
One published claim, traced
Worked example.
Verified live on 5 September 2026.
"I think the products that are doing this, they have a very sharp sense of how well their application is performing, and people don't talk about it, because this is their moat."
Follow it yourself: GET /api/claims/e7240228-46cd-41e6-bb59-0a67d8cb5dca returns the claim, its evidence row, and its references. The claim's assertion, its quote, and the evidence row's quote are all the same string - that identity is what "verbatim" means here.
The row carries transcriptKind: publisher_json - the publisher's own transcript, with speakers already named. That field is on every claim, and it tells you how much the attribution rests on.
Honest boundaries
What we do not claim.
- We verify attribution, not accuracy. Whether they were right is not checked.
- "Verbatim" means against the transcript, not the audio. Publisher mistakes survive.
- The date is the episode's publication date. Not the recording date.
- Timed links are conditional. Only where captions carry timing.
- Absence is shown as absence. No generated answers.
Check it yourself
Where the data is.
Everything on this page is checkable through the documented read-only API. The website and public API currently require no account, API key, subscription or checkout - which describes the product today and is not a promise of a permanent pricing model.
GET https://api.podcasts.highsignal.app/api/stats
GET https://api.podcasts.highsignal.app/api/claims/{id}
GET https://api.podcasts.highsignal.app/api/search?q=agents