embed — 602 chunks, projected to 2-D, coloured by source document
MiniLM-L6 (dim 384) → PCA (SVD, top 2 components) → scatter. Six synthetic
documents, ~100 chunks each. PC1+PC2 explain only 12.5% of the variance -- expected for
random-word paragraphs with no real topical structure to separate; a coherent corpus would
cluster tighter. This is what MiniLM actually did with this exact corpus, not a curated demo.
cloud: GCP Cloud Run, europe-west4, NVIDIA L4 · job: embed-gpu ·
execution: embed-gpu-kkkm9 · date: 2026-09-02 ·
image digest: sha256:2031bac0...4ea6d006d ·
own RESULT payload: 602 chunks, dim 384, 14,942 chunks/min, vectors_saved true ·
projection: PCA via SVD on the centred 602×384 matrix, computed locally from the
uploaded vectors.npy/meta.json (not scored, not a rerun) ·
hero built by the Visuals Lead