4,967 active beliefs · 768-dim embeddings reduced to 3 principal components · colored by domain, sized by authority tier, opacity by confidence
4,967 active beliefs (99.4% embedding_768 coverage — used over the 384-dim column, which only covers 94.5%), reduced from 768 to 3 dimensions via PCA. The first 3 components carry only ~11% of total variance — treat the 3D plot as illustrative, the numbers below as the evidence.
Silhouette (cosine) for domain is 0.001 on the full 768 dimensions and
-0.10 on the 3D projection — near zero or negative both ways, confirming this isn't a
projection artifact. k=15 nearest-neighbor same-domain purity is 56.1% against a 21.4% chance
baseline — some signal above chance, but much weaker than the equivalent check on the source
/domains chunk corpus (75.1% purity there — see that repo's visualisation for the
comparison). authority_tier shows essentially no separation either (silhouette -0.05).
Chunks carry meaningfully more retrievable domain structure than the beliefs later extracted from them. This doesn't yet prove the extractor is responsible — plausible alternative causes include generic belief wording from the prompts, evidence-excerpt-to-belief-statement compression, embedding model/preprocessing differences between the two pipelines, cross-domain aggregation, or formulaic language making otherwise-distinct beliefs converge in embedding space.
The decisive next experiment is a paired comparison: embed the same triple — source chunk →
exact evidence excerpt → composed belief statement — with the same model, and measure cosine
drift by domain, prompt ID/version, and epistemic mode. That would identify where domain
signal disappears along the pipeline, rather than guessing at which stage is responsible. Not yet
run — flagged here as the recommended follow-up, not a finding.