Sentinel Signal

Sentinel Policy M11: Heterogeneous Medicaid Expansion (v1.0.698)

Source: docs/sentinel-policy-m11-heterogeneous-medicaid-expansion-v1.0.698.md

Document Content

Sentinel Policy M11: Heterogeneous Medicaid Expansion (v1.0.698)

What this milestone adds

Per the roadmap's M11 bullet: "3rd source adapter proves the abstraction is genuinely stable." Added TmhpProviderManualAdapter (adapters/tmhp_medicaid.py) targeting the Texas Medicaid Provider Procedures Manual (TMPPM), published monthly by TMHP at tmhp.com -- chosen specifically because it is structurally nothing like CMS's clean REST API (M2): monthly-dated per-chapter PDFs, an HTML (not JSON) Release Notes changelog, and its own distinct terms-of-use language. See docs/adr/sentinel-policy-adr-004-tmhp-state-medicaid-adapter.md for the full design and every real finding along the way.

The actual exit criterion, stated plainly: no new Alembic migration, and no changes to fetch.py, rights/registry.py, ingest.py, policyintel.py, diffengine.py, extraction.py, or any M9/M10 customer-surface code. The one new file (tmhp_medicaid.py) plus one new if branch in registry.py is the entire footprint -- everything downstream of discover()/fetch()/normalize() already worked identically regardless of source, and this milestone proves it rather than needing to build it.

  • Two genuinely new parsing capabilities this repo never needed
  • against CMS's JSON API: HTML table parsing (beautifulsoup4, for the Release Notes page) and PDF text extraction (pypdf, for chapter content). Both added as narrow, well-established dependencies, not hand-rolled parsing.

  • Real terms-of-use evaluated, not assumed: TMHP's own real language
  • restricts use to "personal use... directly participating in [Texas Medicaid]" -- a genuinely not-yet-commercially-cleared source. The registered Source row stays terms_review_state=DISCOVERED, no exception (M0's rule, unchanged by this milestone).

  • Rights registry regression, proven not assumed: TMHP's manual
  • contains real CPT (AMA) and CDT (ADA) content, the same two code systems M0's registry already classifies RED from CMS content. A test proves it's suppressed with zero registry changes.

What testing caught -- real findings, not glossed over

  • A real ordering bug, found by the tests, not by inspection: TMHP's
  • Release Notes page is explicitly newest-first. policyintel.py's version-chaining only links a new version to an existing lower- ordinal one -- it never retroactively relinks. Returning discover() results in the page's own order would leave a from-scratch multi- edition backfill entirely unlinked (no previous_version_id anywhere, so M4's diff engine never runs). Fixed by sorting discover()'s results oldest-edition-first -- a real adapter-design implication any future newest-first source needs too, not a TMHP-only quirk.

  • A route/host mismatch caught immediately by the first test run:
  • real captured Release Notes HTML links to absolute https:// www.tmhp.com/... URLs. discover() originally passed those straight through, which meant base_url overrides (needed for tests to point at a local fixture server, same pattern CMSCoverageAdapter uses) never actually took effect -- every fetch failed the hostname allowlist check. Fixed by extracting only the URL path from parsed HTML and always re-anchoring it to self._base_url, which is also a real security improvement (never blindly trusts an absolute URL parsed out of fetched content).

  • Malformed PDFs degrade, proven not asserted: normalize() runs
  • outside run_discovery()'s per-resource try/except (confirmed by reading ingest.py), so a raised exception there would crash the whole run, not just one resource. A real fixture test with corrupt PDF bytes confirms this degrades to a zero-section PolicyVersion (the same, already-handled "no section for content not provided" case), not a crash.

  • A structural inconsistency within TMHP's own page: the styled
  • "Handbook / Related Articles" header row sometimes lives inside <thead>, sometimes as the first <tbody> row, across different months on the same page. The parser filters by "does this row's first cell link to a .pdf" rather than depending on position -- robust to both variants without needing to special-case either.

  • **A real production concurrency bug, found during live verification
  • itself (v1.0.699 follow-up commit)**: registering the new Texas source and manually triggering discovery within the same minute raced against the scheduler's own poll_once(), which independently picked up the same source (last_attempt_at was still None) and started a second, concurrent run_discovery() against it. Both sessions repeatedly UPDATEd the same sources row for as long as they ran -- harmless for a fast CMS-style run, but TMHP's own real ~8-minute, 81-resource run gave a wide enough window for the two sessions to lock-contend long enough to hit Postgres's statement_timeout and 500. Not TMHP-specific -- this race was always possible for any source, just never exercised by a run slow enough to make it land in practice until now. Fixed in ingest.py::run_discovery() with a concurrent-run guard (skip if another run is already RUNNING for the same source, with a 60-minute staleness escape hatch so a crashed process can't block a source's discovery forever) -- 2 new tests, 205 passing total. See the "Live verification" section below for how this was caught and confirmed fixed against the real site.

Files changed

  • policy/src/sentinel_policy/adapters/tmhp_medicaid.py -- new.
  • policy/src/sentinel_policy/adapters/registry.py -- one new
  • if adapter_type == "tmhp_provider_manual" branch.

  • policy/pyproject.toml -- beautifulsoup4, pypdf.
  • policy/tests/test_tmhp_medicaid_adapter.py -- new, 7 tests against
  • real captured fixtures (a real Release Notes page, two real chapter PDF editions of the same chapter, one month apart).

  • policy/tests/fixtures/tmhp_medicaid/ -- new: `release_notes_2026.
  • html (real, full, unmodified capture), release_notes_2026_small. html (a real two-month trim, same pattern as CMS's own _small fixture), claims_filing_2026-06.pdf / claims_filing_2026-08.pdf` (real captured chapter editions, genuinely different content -- used to prove M3/M4 detect a real month-over-month change).

  • policy/src/sentinel_policy/main.py, policy/tests/test_api.py --
  • milestone string.

Test results

PYTHONPATH=policy/src:policy/tests python -m pytest policy/tests -q -- 205 passing (up from 196: 203 for the adapter itself, +2 for the concurrency-guard fix below).

Live verification (completed, not deferred)

Registered the real tx-medicaid-tmppm source in production and ran the full pipeline against the live TMHP site end to end:

  • Discovery: a real 81-chapter discovery run against tmhp.com
  • succeeded (resources_discovered_count=81, documents_downloaded_count=81, ~8 minutes -- this is also what exposed the concurrency bug above). A second run (post-fix) confirmed real idempotency: 3 total discovery runs (2 manual + 1 scheduler, the latter two racing concurrently pre-fix) produced 69 distinct DocumentAsset rows across 243 DocumentObservations -- content-hash dedup working exactly as designed, no duplicates from the earlier race either.

  • M3/M4: 22 real Policy rows, 69 real PolicyVersions, and real
  • CandidateDiff rows with has_material_change=true from genuine month-over-month content differences (e.g. 1_04_client_eligibility, 2_15_outpatient_drug, 2_13_med_specs_and_phys_srvs) -- all through the unmodified M3/M4 pipeline.

  • M5/M6, a real bounded extraction run: targeted the smallest real
  • chapter with a material change (1_03_electronic_data_interchange, ~16.6K normalized-text chars -- deliberately not the 246K-char claims_filing chapter ADR-004 flags as a real, larger LLM-context/ cost concern). Produced 13 real ExtractedFact/ChangeEvent rows with verified evidence spans against real content (month-banner updates, an SSL/VPN -> SFTP wording change, [Revised] heading tags), all AUTO_PUBLISHED with full confidence components.

  • M9/M10 customer surfaces: confirmed GET /v1/changes/{id},
  • GET /v1/policies/{id}, GET /v1/sources/{id} all 404 for this real internally-visible Texas content; GET /v1/changes, /v1/changes.csv, and the list_changes MCP tool (all filtered to this policy) each correctly return empty -- the exact same fail-closed behavior every other DISCOVERED source gets, no special-casing anywhere in the customer-facing layer.

No commercial/customer-visible surface shows any Texas content -- this milestone proves the pipeline works end to end for a genuinely heterogeneous source, not a second commercially-cleared one.