Keyword Research in the AI Era: What Still Matters
Keyword research in the AI era means reading AI Overviews and SERPs alongside search volume, not instead of it. Here's the workflow, with real numbers.
A founder we'll call the usual case opens a keyword tool, types in something like "customer onboarding checklist," and finds eleven thousand monthly searches. That number feels like a green light. It is not a brief. It's a founder starting from a keyword and discovering, a few hours later, that the number tells you almost nothing about what to actually write.
That gap between "people search this" and "here's what to publish" is where most keyword research quietly falls apart. The phrase confirms an audience exists. It doesn't say which of the nine already-ranking pages is winning, what they're missing, or why a tenth page would earn a click, a citation, or a read. Getting from the number to the page requires translating the phrase into a real question, then finding an angle nobody else has bothered to take. That's the actual work of keyword research in the AI era, and it looks different from the workflow most of us learned five years ago.
This piece draws on operational experience building gap-finding and content-analysis tools for a working SEO and GEO practice, not on academic research. Where a claim needs a source outside that experience, it's cited directly.
A keyword list is not a content plan
A spreadsheet of keywords is evidence, not a plan, and treating it like one is where most content programs waste their first month. Volume tells you people are looking. It says nothing about what they need once they arrive, which competitor is currently satisfying them well enough, or where the actual opening sits.
Go back to "customer onboarding checklist." Eleven thousand searches a month sounds like a green light for a generic listicle. Pull up the top ten results and you'll usually find eight are exactly that: a checklist, mildly reordered, with stock advice about welcome emails and product tours. The keyword doesn't tell you this. Reading the field does. Somewhere in that same cluster is a narrower, less obvious question, maybe "how do you sequence onboarding for a product that requires a data import before it's useful," and that question has almost no good answers. The keyword tool never surfaces that distinction on its own. It takes a person looking at the results and asking what's actually missing.
This is also why a keyword-first workflow tends to produce content that reads like it was written to hit a phrase rather than answer a person. Search engines and answer engines are both increasingly built to discount exactly that. Google's own guidance on helpful content warns against pages built primarily to attract search traffic rather than serve a reader, and against mass-produced material assembled because a topic is trending rather than because someone had something to say about it [source: https://developers.google.com/search/docs/fundamentals/creating-helpful-content]. A keyword list can point you at a topic. It can't tell you what to say about it.
What search data still tells you
Keyword data still does three things nothing else does as well. It proves demand exists. It shows you the exact words your audience uses. It reveals what someone is trying to do when they type a phrase in. Keep it for that. Stop asking it to write the brief.
A quick way to see all three at once, using the onboarding example:
- Demand: 11,000 monthly searches confirms the topic is worth an hour of your time, before it's worth a week.
- Language: the audience says "onboarding," "activation," or "getting new users to stick," and each phrasing signals a slightly different reader.
- Intent: "what is customer onboarding" is someone learning, "customer onboarding software" is someone about to buy, "customer onboarding checklist" is someone about to do the work themselves.
Demand is the most basic and most underrated of the three. It stops you from spending a week on a topic nobody is searching for, which happens more often than founders like to admit, usually because the topic feels important internally without being something a customer would ever type into a search bar. Volume is a check on your own instincts, and your instincts are wrong in both directions more often than you'd expect.
Where this breaks is the moment volume gets treated as opportunity. A keyword can have real demand and be entirely saturated with strong, nearly identical answers already. Demand tells you people want something. It never tells you there's room for you to be the one who gives it to them. That second question needs a different kind of research, closer to competitive reading than keyword pulling, and it's covered in more depth in how to research a topic so it actually differentiates.
How AI changes the research job
Keyword research in the AI era hasn't shrunk. It's added a step. On top of the keyword tool, you now query the answer engines directly, the same way your reader would, and read what comes back with the same scrutiny you'd apply to a competitor's article.
The evidence for the shift. Pew Research found that when Google shows an AI summary, users click a traditional result in only 8% of visits, against 15% when no summary appears. They click a link inside the summary itself just 1% of the time [source: https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/]. Semrush tracked AI Overviews appearing on 6.49% of queries in January 2025. That climbed to 24.61% by July, and it was still showing on 15.69% in November, spreading well beyond informational searches into commercial and transactional ones too [source: https://www.semrush.com/blog/semrush-ai-overviews-study/]. A separate Semrush study on commercial intent found AI Overviews grew 71% across commercial search results over six months, appearing more on commercial queries than transactional ones [source: https://www.semrush.com/blog/ai-overviews-commercial-search-study/]. Being the tenth link on a results page was never a great position. It's a worse one now.
The workflow this creates. Run the same query through three lenses and the differences are clear.
| Source | What it gives you | What it misses |
|---|---|---|
| Keyword tool | A phrase and a number, demand roughly worded | No sense of what's actually being asked |
| AI Overview | A synthesized answer, usually competent | Nuance, thinner than the best individual page it drew from |
| Manual SERP review | Which page is winning on topical authority, where the argument is missing | Takes longer, no volume metric attached |
None of the three replaces the others. The keyword tool tells you people are asking. The AI Overview tells you what's currently being said. Reading the results yourself tells you what's still worth saying.
This is also where entity-based search and information gain start to matter more than they used to. Answer engines aren't just matching a string, they're building a picture of the entities connected to a topic, the tools, the roles, the steps, and looking for pages that add something to that picture rather than repeating it. Topical authority, in this context, isn't a vague reputation score. It's a page's ability to be more complete on the entities and sub-questions around a topic than what already exists, which is a different bar than matching the keyword well.
That bar also shifts depending on what the searcher is actually trying to do. An informational query like "what is customer onboarding" rewards depth and clear explanation, since the reader is still forming the question. A commercial query like "best customer onboarding software" rewards comparison and evidence of having actually used the options, since AI Overviews on commercial queries tend to lean on named tools and features. A transactional query, closer to "customer onboarding software pricing," rewards specificity over synthesis, since there's less for an AI summary to usefully compress. Treating all three the same way is a common shortcut, and it's usually visible in the output.
Google itself has said AI Overviews are meant to surface a wider range of sources so people can click out and explore further, not to replace the web entirely [source: https://developers.google.com/search/blog/2025/05/succeeding-in-ai-search?hl=en]. That's an opening, but only for a page built to be the clearer, more complete source, one with enough topical authority and citation-worthy specifics that an answer engine has a reason to point to it. Structuring content so an answer engine has a reason to cite it is its own discipline, covered in generative engine optimization.
None of this replaces judgment. An engine can tell you what the current synthesis looks like. It can't tell you whether the gap is worth filling, or whether your angle is actually different from the eight other pages that also noticed the same gap.
Building a topic around a real gap
A keyword becomes useful the moment you can say what your page adds that the current field doesn't already have. Without that answer, you're writing the ninth version of something that already exists nine times.
A workable framework for this, based on what we've built into our own tooling, has three checks, and a topic should clear all three before it gets written.
- Novelty against your own archive. Does the topic already exist on your site under a different title? Our competitor-gap engine deduplicates suggested topics against everything already published using embedding similarity, and we set that threshold deliberately high, because early versions kept proposing articles the client had already written under a slightly different title. That's an internal observation from building the tool, not a benchmark anyone else has verified, and it's worth being clear about that distinction.
- Novelty against the field. Does the topic already exist, well answered, on the top-ranking pages and in the AI Overview? If six of the top ten results already say the same thing, the gap isn't the topic, it's whatever those six pages left out.
- Fit against the brand. Does the angle match who you actually are? We filter opportunity suggestions by brand fit on purpose, which means a premium, narrowly positioned brand gets fewer suggestions, not more, and they're deliberately specific rather than broad.
Take a concrete case. A client in workplace software kept surfacing "employee onboarding template" as a gap, volume in the low thousands, seemingly untouched. Pulling the actual SERP and the AI Overview readout showed the term already answered thoroughly, six template roundups in the top ten, an AI summary that named the same three tools every time. The missing angle wasn't the template. It was what happens after week one, when the template stops being useful and nobody's written about the handoff from HR to a manager. That became the actual topic, not a checklist, but the specific failure point checklists don't cover.
Fit matters as much as gaps do. One user read our tool's narrow suggestion list as the tool being broken, when it was actually the tool refusing to suggest topics that didn't match the brand it was built for. If a system is going to be opinionated about what's worth writing, it has to say so plainly, or the opinion just looks like a bug.
Precision has to run through the whole pipeline, not just the topic suggestions. Our SEO analyzer used to confidently report tables and FAQ sections in articles that had neither. It was grading from impressions rather than evidence. The fix wasn't a cleverer prompt, it was ground truth: we count headings, tables, images, and lists with regex first, then hand those exact numbers to the model as facts it isn't allowed to contradict. The false positives stopped immediately. We apply the same discipline to numbers in a draft. A statistic gets used only if it traces back to the brief, the source material, or research pulled from an actual web search with a URL attached. An uncited number is treated as invented, because when we've spot-checked them internally, the uncited ones usually were.
That's the standard a gap has to clear too. A real one is specific enough to name, and the page you build around it has to hold up under the same scrutiny you'd apply to a source you were citing.
Less work, more on-brand content
Austen runs this whole workflow for you: from research to on-brand drafts that get found by Google and AI.
Start freeMore in Research & Differentiation
-
Research & Differentiation: How to Write Content That Adds to the Conversation
Learn what content research that differentiates actually looks like: the five sources of real differentiation, how to spot a content gap, and why originality wins in search and AI answers.
-
How to Use Original Research to Stand Out (and Get Cited)
What makes original research content citeable, how to package a finding with method, sample, and caveats, and when a research system should refuse to publish a claim.
-
How to Find Content Gaps Your Competitors Left
A practical method for finding content gaps, built for small B2B teams: how to separate topic gaps from intent gaps, and turn either into a brief that ranks.