Semantic keyword clustering in Python is a short script that turns every keyword into an embedding (a list of numbers that captures meaning) with the sentence-transformers library, then groups those vectors with a clustering algorithm from scikit-learn. Keywords that mean the same thing land in the same group, even when they share no words. The whole semantic keyword clustering Python workflow below is about 50 lines, runs on a normal laptop CPU and costs nothing.
If you want the concept first, meaning why grouping by meaning beats grouping by shared words, read my comparison of keyword clustering vs semantic clustering. This post is only the code and the decisions around it.
What Do You Need Before You Run the Code?
You need Python 3.10 or newer, three libraries and a CSV of keywords. That’s it. No API key, no paid tool.
Install the libraries with pip. The first time you load a model, sentence-transformers downloads it from Hugging Face and caches it, so the first run is slower than every run after.
pip install -U sentence-transformers scikit-learn pandas
Every block in this post was run end to end on September 29, 2026 with sentence-transformers 6.1, scikit-learn 1.7, pandas 2.3 and PyTorch 2.14. The parameters match the current scikit-learn AgglomerativeClustering docs. Older tutorials pass affinity="cosine", which newer scikit-learn versions reject. The parameter is now called metric.
Your input file needs one column called keyword. Here is the sample list I use in this tutorial, 31 rows from a made-up home coffee site, with two deliberate duplicates at the bottom:
keyword
best espresso machine
espresso machine for beginners
best home espresso maker
top rated espresso machines
espresso machine under 500
cheap espresso machine
budget espresso maker
how to descale an espresso machine
descaling espresso machine with vinegar
how often to descale espresso machine
best coffee grinder
burr grinder vs blade grinder
best burr grinder for espresso
conical burr grinder
what grind size for espresso
espresso grind size chart
how fine to grind coffee for espresso
how to make espresso without a machine
espresso at home without machine
moka pot espresso
how to use a moka pot
moka pot coffee ratio
latte vs cappuccino
difference between latte and flat white
cappuccino vs flat white
how to froth milk without a frother
milk frother for latte
best milk for latte art
Best Espresso Machine
best espresso machines
Save it as keywords.csv next to the script.
Four Questions That Change the Code
Before you copy anything, answer these. Each one changes a line or two of the script.
- How many keywords do you have? Under about 10,000, agglomerative clustering (the main path below) is simple and accurate. Well above that, switch to the fast community detection branch.
- What language are they in? English lists work with
all-MiniLM-L6-v2. Other languages need a multilingual model, and not every multilingual model covers every language. - Do you already know how many pages you want? Almost nobody does. If you genuinely do, K-means with a fixed cluster count is an option. If not, use a distance threshold and let the data decide.
- Will clusters become pages without review? If yes, add the optional results-page check at the end. Meaning alone can’t tell you when two close queries need separate pages.
The Full Semantic Keyword Clustering Python Script
This is the main path: English keywords, a list of a few hundred to a few thousand rows, and no fixed cluster count. Read the walkthrough under the code before you run it on real data.
import sys
import numpy as np
import pandas as pd
from sentence_transformers import SentenceTransformer
from sklearn.cluster import AgglomerativeClustering
INPUT_FILE = "keywords.csv"
OUTPUT_FILE = "keyword_clusters.csv"
MODEL_NAME = "all-MiniLM-L6-v2"
DISTANCE_THRESHOLD = float(sys.argv[1]) if len(sys.argv) > 1 else 0.4
df = pd.read_csv(INPUT_FILE)
df["keyword"] = df["keyword"].astype(str).str.lower().str.strip()
df["keyword"] = df["keyword"].str.replace(r"\s+", " ", regex=True)
df = df[df["keyword"] != ""].drop_duplicates(subset="keyword").reset_index(drop=True)
keywords = df["keyword"].tolist()
print(f"{len(keywords)} unique keywords")
model = SentenceTransformer(MODEL_NAME)
embeddings = model.encode(
keywords,
batch_size=64,
show_progress_bar=True,
normalize_embeddings=True,
)
clusterer = AgglomerativeClustering(
n_clusters=None,
metric="cosine",
linkage="average",
distance_threshold=DISTANCE_THRESHOLD,
)
df["cluster_id"] = clusterer.fit_predict(embeddings)
labels = {}
for cluster_id, idx in df.groupby("cluster_id").groups.items():
idx = list(idx)
centroid = embeddings[idx].mean(axis=0)
closest = idx[int(np.argmax(embeddings[idx] @ centroid))]
labels[cluster_id] = df.loc[closest, "keyword"]
df["cluster_label"] = df["cluster_id"].map(labels)
df["cluster_size"] = df.groupby("cluster_id")["keyword"].transform("size")
df = df.sort_values(["cluster_size", "cluster_label"], ascending=[False, True])
df.to_csv(OUTPUT_FILE, index=False)
print(df["cluster_id"].nunique(), "clusters saved to", OUTPUT_FILE)
for label, group in df.groupby("cluster_label", sort=False):
print(f"\n[{label}]")
for kw in group["keyword"]:
print(" -", kw)
Run it with python cluster_keywords.py for the default threshold, or pass a number such as python cluster_keywords.py 0.3 to try another one.
Step 1: Clean Before You Embed
The cleaning step is boring and it matters. Lowercasing, trimming spaces and dropping exact duplicates turned the 31 sample rows into 29 unique keywords, because “Best Espresso Machine” and a copy with a trailing space were the same query.
Notice what it did not merge. “best espresso machine” and “best espresso machines” both survive, because they’re different strings. The embedding step handles that pair later, so don’t write clever plural rules here.
Step 2: Turn Keywords Into Embeddings
all-MiniLM-L6-v2 is my default for English keyword lists. The sentence-transformers team describes it as about 5 times faster than their best-quality general model, all-mpnet-base-v2, while still offering good quality. For short queries of three to eight words, I start with the small model and only switch if the groups look wrong.
The normalize_embeddings=True flag scales every vector to length 1. That makes a plain dot product equal to cosine similarity, which the labelling step relies on. Keep it on.
Step 3: Cluster With a Distance Threshold
AgglomerativeClustering starts with every keyword as its own cluster and keeps merging the closest pair until the next merge would cross your threshold. With metric="cosine", distance is 1 minus similarity, so a threshold of 0.4 means “merge groups whose average similarity is at least about 0.6.”
Two rules trip people up. You must set n_clusters=None when you set a threshold, and the ward linkage only works with Euclidean distance. Use average for keywords. In my experience it gives tidier groups than complete, which splits too eagerly, or single, which chains unrelated keywords together.
Step 4: Name Each Cluster and Export
Each cluster is named after the keyword closest to its centre, not the first or the longest one. That tends to pick the most typical phrasing, which is usually the best working title for the page. If you also have search volume, add it as a column and sort by it inside each cluster.
Here is what the script prints for the coffee list at the default threshold:
29 unique keywords
8 clusters saved to keyword_clusters.csv
[best espresso machine]
- best espresso machine
- espresso machine for beginners
- best home espresso maker
- top rated espresso machines
- espresso machine under 500
- cheap espresso machine
- budget espresso maker
- how to make espresso without a machine
- espresso at home without machine
- best espresso machines
[best burr grinder for espresso]
- best coffee grinder
- burr grinder vs blade grinder
- best burr grinder for espresso
- conical burr grinder
[how to descale an espresso machine]
- how to descale an espresso machine
- descaling espresso machine with vinegar
- how often to descale espresso machine
[milk frother for latte]
- how to froth milk without a frother
- milk frother for latte
- best milk for latte art
[moka pot espresso]
- moka pot espresso
- how to use a moka pot
- moka pot coffee ratio
[what grind size for espresso]
- what grind size for espresso
- espresso grind size chart
- how fine to grind coffee for espresso
[latte vs cappuccino]
- latte vs cappuccino
- cappuccino vs flat white
[difference between latte and flat white]
- difference between latte and flat white
Most of that is right. Descaling, moka pots and grind size are clean, page-ready groups, and the model paired “best espresso machine” with “best home espresso maker” without a single shared word beyond “espresso.”
Now look at the biggest cluster. “How to make espresso without a machine” sits inside a buying cluster, because it scored 0.72 similarity with “best espresso machine.” The words overlap, the intent doesn’t. Somebody searching it wants a method, not a product. That is the single most common mistake embeddings make on keyword lists, and it’s why I never publish clusters without a read-through.
How Do You Choose the Right Distance Threshold?
Don’t guess it. Sweep it. This short loop reuses the embeddings from the main script and prints how the groups change as the threshold moves:
for threshold in [0.2, 0.3, 0.4, 0.5, 0.6]:
labels = AgglomerativeClustering(
n_clusters=None,
metric="cosine",
linkage="average",
distance_threshold=threshold,
).fit_predict(embeddings)
sizes = np.bincount(labels)
singletons = int((sizes == 1).sum())
print(f"threshold {threshold}: {len(sizes)} clusters, "
f"{singletons} single-keyword clusters, largest has {sizes.max()}")
On the coffee list it printed this:
| Threshold | Clusters | Single-Keyword Clusters | Largest Cluster |
|---|---|---|---|
| 0.2 | 20 | 14 | 5 |
| 0.3 | 14 | 8 | 10 |
| 0.4 | 8 | 1 | 10 |
| 0.5 | 5 | 0 | 13 |
| 0.6 | 4 | 0 | 20 |
You’ll see the same pattern on most lists. Too low, and half your keywords sit alone, which means one page per query and a pile of near-duplicates. Too high, and one giant cluster swallows three or four different pages.
My rule of thumb: pick the lowest threshold where singletons drop below roughly 10% of the list, then read the three largest clusters by hand. On the coffee list that’s 0.4. Treat any number, including mine, as a starting point for your data, not a constant.
A threshold also can’t fix a pair that sits right on the line. “Latte vs cappuccino” and “difference between latte and flat white” scored 0.58 similarity, so at 0.4 they split.
Both are comparison queries about milk drinks, and I’d happily answer them on one page. The script gets you most of the way. You make the last call.
Branch: What If You Have Tens of Thousands of Keywords?
Agglomerative clustering compares every keyword with every other keyword. For 10,000 keywords that’s roughly 50 million pairwise distances, about 400 MB of memory at 64-bit precision, and the number grows with the square of the list. At 50,000 rows it’s about 1.25 billion distances and roughly 10 GB, which stops being practical on a laptop.
For big lists, swap Step 3 for the community detection function that ships with sentence-transformers. The sentence-transformers clustering docs describe it as fast enough to cluster 50,000 sentences in under 5 seconds on their hardware.
from sentence_transformers import util
communities = util.community_detection(
embeddings,
threshold=0.75,
min_community_size=2,
batch_size=1024,
)
cluster_id = np.full(len(keywords), -1)
for i, members in enumerate(communities):
cluster_id[members] = i
next_id = len(communities)
for i in np.where(cluster_id == -1)[0]:
cluster_id[i] = next_id
next_id += 1
df["cluster_id"] = cluster_id
Two differences matter. Here the threshold is a similarity (0.75 is the library default), not a distance, so higher means stricter. And any keyword that doesn’t fit a community is simply left out of the result, which is why the second loop gives each leftover its own cluster.
Skip that loop and you silently lose keywords. On the coffee list, community detection at 0.75 placed only 19 of the 29 keywords into groups.
After this block, run Step 4 from the main script unchanged.
Branch: What About Non-English Keywords?
Swap the model name and keep everything else. paraphrase-multilingual-MiniLM-L12-v2 covers 50+ languages, but check its Hugging Face model card for yours before you trust it. Bengali, for example, isn’t in that model’s listed languages.
For languages like Bengali, intfloat/multilingual-e5-small does list it. That model family expects a prefix on every input, so encode ["query: " + k for k in keywords] instead of the raw list. Then test on a sample of 50 keywords you know well before you run the full list.
Optional: Check Close Pairs Against the Live Results
Embeddings know what words mean. They don’t know what Google ranks. The fix for cases like the espresso-without-a-machine query is to compare the top 10 URLs for keyword pairs inside the same cluster, and flag pairs whose results barely overlap.
This works on a CSV you already have from a rank tracker or SERP export, with the columns keyword, position and url. It doesn’t scrape anything:
import itertools
serps = pd.read_csv("serps.csv")
top10 = serps[serps["position"] <= 10].groupby("keyword")["url"].apply(set)
def serp_overlap(a, b):
if a not in top10 or b not in top10:
return None
return len(top10[a] & top10[b]) / 10
flags = []
for label, group in df.groupby("cluster_label"):
for a, b in itertools.combinations(group["keyword"], 2):
share = serp_overlap(a, b)
if share is not None and share < 0.3:
flags.append((label, a, b, share))
pd.DataFrame(
flags, columns=["cluster", "keyword_a", "keyword_b", "serp_overlap"]
).to_csv("split_candidates.csv", index=False)
print(len(flags), "pairs to review")
The 0.3 cutoff means fewer than 3 shared URLs out of 10, the same threshold I unpack in my walkthrough of SERP based keyword clustering. Every flagged pair is a candidate to split into its own page. Use it as a review list, not an automatic rule, and only pay for SERP data on the clusters you plan to publish.
Decision Matrix: Which Version of the Script Do You Need?
| If Your List Is | And You Need | Then Use |
|---|---|---|
| Under 10,000 English keywords | Page-level clusters | The main script, threshold swept from 0.3 to 0.45 |
| Over 20,000 keywords | Topic-level groups fast | util.community_detection plus the leftover loop |
| Any size, non-English | Either | A multilingual model that lists your language, tested on 50 rows first |
| A fixed number of pages already decided | Exactly N groups | KMeans(n_clusters=N) on the same embeddings |
| Clusters going straight into a content plan | Intent-safe groups | Main script plus the results-page overlap check |
What Can’t the Script Decide for You?
It can’t decide intent, priority or ownership. A cluster tells you which queries belong together. It doesn’t tell you which cluster is worth a page, which one an existing URL already owns, or which one you can realistically win.
That’s the part of the job I still do by hand. When I take a clustered list into planning, I map each cluster to one URL and check it against the site’s existing pages, which is exactly what keyword mapping covers. If you’d rather not write code at all, I compared the paid and free options in keyword clustering tools.
And if you’d prefer someone else to handle it end to end, Skyranko’s keyword research service delivers intent-checked clusters mapped to pages, with 5 keywords free to try.
Frequently Asked Questions
Is Semantic Keyword Clustering in Python Free?
Yes. sentence-transformers, scikit-learn and pandas are open source, and the models download free from Hugging Face. The only cost is your time and, if you add the optional overlap check, whatever you pay for SERP data.
Do I Need a GPU to Cluster Keywords With Python?
No. Keywords are short, so a laptop CPU embeds a few thousand of them with all-MiniLM-L6-v2 quickly. A GPU helps once you’re into tens of thousands of rows or switch to a larger model.
Why Not Use OpenAI Embeddings Instead?
You can. Replace Step 2 with calls to any embedding API and keep Steps 3 and 4. I start with a local model because it’s free, private and repeatable, and on short queries the grouping difference is usually small.
Does This Replace Checking the Search Results?
No. This script groups by meaning, while the results page shows how Google actually treats each query. Tools that group by shared ranking URLs charge per keyword for that view. The optional overlap block above gives you a small-scale version of it for the clusters you care about.
Last updated: September 2026 by Mizanur Rahman



