Practice — Clustering & Dimensionality Reduction (6 questions)
k-means Local Minimum by Hand Permalink →
You have nine 1-D points: 1, 2, 3, 10, 11, 12, 20, 21, 22 and want
k = 3 clusters. A colleague initialises the centroids at 1, 2, 3
and runs Lloyd's algorithm to convergence.
- Run the algorithm by hand (assign, then update, repeat). What are the final centroids and clusters, and what is the inertia J?
- What is the inertia of the "obvious" clustering
{1,2,3}, {10,11,12}, {20,21,22}? What does the comparison tell you? - How would k-means++ initialisation have made the bad outcome unlikely?
Share this question
Upgrading to a Bigger Embedding Made Search Worse Permalink →
Your k-NN recommender uses 2-dimensional item embeddings and works well: for a typical query point, the nearest neighbor is clearly much closer than the farthest point in the set. You swap in a fancier 1,000-dimensional embedding model, expecting sharper, more meaningful neighbors. Instead, recommendations get noticeably worse — nearest and farthest neighbors start looking almost interchangeable.
As the number of (roughly independent, noise-like) dimensions grows very large, what happens to the ratio of the farthest point's distance to the nearest point's distance, for a fixed query?
Share this question