← Cloud Olympics catalog

embed — 602 chunks, projected to 2-D, coloured by source document

MiniLM-L6 (dim 384) → PCA (SVD, top 2 components) → scatter. Six synthetic documents, ~100 chunks each. PC1+PC2 explain only 12.5% of the variance -- expected for random-word paragraphs with no real topical structure to separate; a coherent corpus would cluster tighter. This is what MiniLM actually did with this exact corpus, not a curated demo.

synthetic_00.mdsynthetic_01.mdsynthetic_02.mdsynthetic_03.mdsynthetic_04.mdsynthetic_05.md PC1 (6.7% var) PC2 (5.8% var)
cloud: GCP Cloud Run, europe-west4, NVIDIA L4 · job: embed-gpu · execution: embed-gpu-kkkm9 · date: 2026-09-02 · image digest: sha256:2031bac0...4ea6d006d · own RESULT payload: 602 chunks, dim 384, 14,942 chunks/min, vectors_saved true · projection: PCA via SVD on the centred 602×384 matrix, computed locally from the uploaded vectors.npy/meta.json (not scored, not a rerun) · hero built by the Visuals Lead