Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Clustering & Dimensionality Reduction (6 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

k-means Local Minimum by Hand Permalink →

You have nine 1-D points: 1, 2, 3, 10, 11, 12, 20, 21, 22 and want k = 3 clusters. A colleague initialises the centroids at 1, 2, 3 and runs Lloyd's algorithm to convergence.

  1. Run the algorithm by hand (assign, then update, repeat). What are the final centroids and clusters, and what is the inertia J?
  2. What is the inertia of the "obvious" clustering {1,2,3}, {10,11,12}, {20,21,22}? What does the comparison tell you?
  3. How would k-means++ initialisation have made the bad outcome unlikely?

Share this question

Intermediate Open Pro

Computing and Interpreting Silhouette Scores

Unlock this question →
Intermediate Open Pro

Choosing Between k-means, DBSCAN and GMM

Unlock this question →
Intermediate Open Pro

PCA on a Small Dataset

Unlock this question →
Intermediate Open Pro

Misreading a t-SNE Plot

Unlock this question →
Advanced Open Free

Upgrading to a Bigger Embedding Made Search Worse Permalink →

Your k-NN recommender uses 2-dimensional item embeddings and works well: for a typical query point, the nearest neighbor is clearly much closer than the farthest point in the set. You swap in a fancier 1,000-dimensional embedding model, expecting sharper, more meaningful neighbors. Instead, recommendations get noticeably worse — nearest and farthest neighbors start looking almost interchangeable.

As the number of (roughly independent, noise-like) dimensions grows very large, what happens to the ratio of the farthest point's distance to the nearest point's distance, for a fixed query?

Share this question

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.