Loading…
Each point is one chunk, stratified-sampled (capped 500/domain so small domains stay visible) from the 106,104-chunk register, embedded with nomic-embed-text:v1.5 (768d, unit-normalized) and reduced to 3 axes with PCA. The first 3 components carry only 16.5% of total variance — this 3D view is a lossy projection. Numbers below are computed on the full 768 dimensions, not read off the picture; treat the plot as illustrative and the numbers as the actual evidence.
Red rings mark chunks with a known quality signal: chunk authority_tier ≠ its document's current tier (corpus-wide drift, one of the two axes this repair programme tracks), and a heuristic character-duplication proxy for mojibake/OCR corruption (not a verified list — a text-corruption density score, thresholded; see analyze_domains_chunks.py's docstring for the spot-check methodology).
Domains separate locally more than the global picture suggests. k=15 nearest-neighbor
same-domain purity is 75.1% against a 16.7% chance baseline — local neighborhoods are strongly
domain-consistent. construction.nhbc is the most distinct (97.4% purity), a narrow
UK-housebuilding vocabulary. construction.geotechnical is the least separated (53.7%),
consistent with its heavy overlap with construction.structural via shared Eurocode
material (EN 1992/1993/1997).
Label-only defects (tier drift) are invisible in embedding space, as they should be. Tier-drift chunks sit at essentially the same cosine-to-centroid as clean chunks (0.790 vs 0.780) — no spatial signature, because only a metadata field is wrong, not the text.
Text-corruption (mojibake) chunks show up as measurable outliers. In every domain with
enough flagged chunks to compare, mojibake-flagged chunks sit further from their domain centroid than
clean chunks in the same domain — most clearly in construction.3dprinting (0.72 vs 0.79)
and construction.geotechnical (0.70 vs 0.78). This corruption pattern is not confined to
one domain — it appears across at least 4 of the 6 sampled categories.