Skip to content
Source triage for keyword research

Rank the sources.See the gaps.Rewrite the query.

LLM Keyword Ranker turns a keyword into a scored reading list: how well each source answers the question you actually have, how much weight it can carry, and which parts of that question nothing in the list answers yet.

Free plan, 25 ranked searches a month, no card.

Two scores
Relevance and trustworthiness, kept apart and never averaged into one comfortable number.
Typed gaps
Five kinds of hole in a result set, each named, each with a different fix.
Plain queries
Refinements written as searches you could have typed yourself, with the reasoning attached.

A worked example throughout this site: the keyword sleep apnea, asked by someone who wants to know whether treatment improves daytime function in adults.

S01Five stages, in order

What happens between typing a keyword and knowing what to read

Read the full walkthrough
  1. 01

    Type the keyword, then say what you actually want to know

    A keyword is a handle, not a question. The first screen asks for both, because everything downstream is scored against the question rather than the string.

  2. 02

    Read a list ordered by two scores, never one

    Every source comes back with a relevance score and a trust score, side by side. They are never averaged, because a compelling blog post and a rigorous paper about the wrong question are different problems.

  3. 03

    Open any score and see what produced it

    A score you cannot interrogate is a number you should not trust. Click either bar and the six trust signals open up, each with the evidence that moved it.

  4. 04

    See what your results never answer

    The gap map is the part that changes how people work. It reads your question against the whole result set and marks the sub-questions nothing in it addresses.

  5. 05

    Take a rewritten query, or ignore it

    Every gap produces a query that would close it. Each one arrives with the change it makes and the reason it makes it, so you can judge the suggestion instead of just running it.

S02What the two scores mean

Trustworthiness and relevance, in plain words

Two different questions get asked about every source, and most search tools only ever ask one of them.

Relevance asks: is this about my question?

Not whether the page contains your keyword. Whether it addresses the thing you want to know. A page can use the words sleep apnea forty times and still be a shopping guide rather than an answer to whether treatment improves how you feel at three in the afternoon. So relevance scores four things: whether the source engages the question behind the keyword, how specific it is to that question rather than to the whole field, whether it studies the population and setting you asked about, and whether your topic is the subject of the piece or a passing mention inside a paragraph about something else.

Trustworthiness asks: how much weight can this carry?

It is not a verdict on honesty and it is not a popularity contest. It is an estimate of how much a claim can bear, built from things you could check yourself if you had the afternoon free: who published it and whether a named person stands behind it, whether it cites primary evidence and whether those citations say what the source claims they say, whether anyone examined it before publication, how old the claim is relative to how fast the field moves, who paid for the work, and whether independent sources reach the same conclusion.

The two are never blended

A highly relevant blog post and a rock-solid paper about an adjacent question are two different problems, and averaging them into one number hides which one you have. A row reading 89 relevance and 24 trust is telling you something worth knowing: this is exactly your question, and the answer on offer is not one you can lean on. You always see both numbers, and you can always open either one to see what produced it.

Clicking a trust score opens this. Every signal shows the material that moved it, and anything that could not be assessed is removed from the total rather than quietly scored as average.

The six trust signals

T1Provenance

Who is behind this?

A named author, a real institution, an editorial address you could write to, a publisher that puts its name on its corrections. Anonymous aggregator pages score low here even on the occasions when they happen to be right.

T2Evidence

What is it built on?

Whether the source cites primary research or somebody else's summary of it, whether those citations exist, and whether they support the sentence they are attached to. Confident, uncited writing loses most of its score right here.

T3Scrutiny

Did anyone check it before you did?

Peer review, editorial review, a preprint that was later published, a version history with visible corrections. A source that has been corrected scores higher than one nobody has ever examined.

T4Currency

Has the field moved since?

Age measured against how fast the topic changes, not against the calendar. A 1998 paper on a measurement technique can be perfectly current; a 1998 paper on treatment guidance is not. We date the claim, not the page.

T5Independence

Who paid, and what did they want?

Declared funding, author affiliations, and whether the publisher sells the thing being evaluated. Rarely disqualifying on its own, always something you should know before you cite it.

T6Corroboration

Does anything else agree?

Whether independent sources reach the same finding, and whether any credible source reaches the opposite one. Contradiction lowers confidence and is reported, rather than being quietly dropped from the list.

The four relevance signals

R1Question match

Does it address what I asked?

Whether the source engages with the question behind your keyword, rather than simply containing the word. This is the signal that separates an answer from a mention.

R2Specificity

Is it about my question or the whole field?

A survey of an entire discipline and a study of your exact sub-question are both relevant, but not equally. Broad overviews are marked as such so they stop crowding out the narrow work.

R3Context match

Is it about my population and setting?

Age group, geography, scale, discipline, time period. Evidence about a different group is still evidence, but it is not evidence about your group, and the difference is scored rather than glossed over.

R4Centrality

Is my keyword the subject or an aside?

Whether the source is about your topic or mentions it once in a paragraph about something else. Passing mentions are the single largest source of noise in ordinary keyword search.

S03Key features

Six things it does that ordinary keyword search does not

All eight features in detail
  • Dual source ranking

    Two scores, never averaged

    Every source is scored twice: how well it answers your question, and how much weight it can carry. You always see both numbers.

    How dual source ranking works
  • Reliability insights with receipts

    Six trust signals, each openable

    Click any trust score and it decomposes into provenance, evidence, scrutiny, currency, independence and corroboration, each with the material behind it.

    How reliability insights with receipts works
  • Relevance past keyword match

    What it is about, not what it contains

    Four relevance signals separate sources that address your question from sources that merely use your words.

    How relevance past keyword match works
  • Knowledge gap detection

    What your results never answer

    The gap map breaks your question into parts and reports the parts nothing in your results addresses, typed by what kind of gap it is.

    How knowledge gap detection works
  • Actionable query refinement

    Rewritten searches with the reasoning attached

    Each gap produces a query that would close it, written in plain search syntax, with the move it makes and why it is worth making.

    How actionable query refinement works
  • Export and share

    Take the work out intact

    Exports carry the scores, the signals, the gaps and your overrides, in the formats reference managers and colleagues actually read.

    How export and share works
S04Knowledge gap insights

The most dangerous result set is the one that feels complete

Ten strong sources will happily let you believe your question is answered. The gap map breaks the question into parts and tells you which parts nothing in your results actually addresses.

Five kinds of gap, because five different fixes

G1Coverage gap
Nothing in your results addresses this part of the question at all. The fix is a new query, and the refinement queue will have written one.
G2Evidence gap
The topic is covered, but only by opinion, anecdote or somebody summarising somebody else. There is writing about it and no research behind the writing.
G3Currency gap
Every source that addresses this predates the most recent change in the field. The answers may still be right, but nothing in your list has tested them since.
G4Population gap
The evidence exists and is sound, and it is about a different group, place or scale than the one you asked about. The most common gap, and the easiest one to read past.
G5Disagreement
Two or more credible sources reach opposite conclusions and nothing in your results resolves it. Flagged as unsettled rather than silently ranked in favour of one side.

Every gap is a claim about your result set, so every gap is checkable. Open one and it lists the sources it examined and why each one fell short.

What the map found for sleep apnea

  • Population gapG4

    Adults over 70

    Every trial in your results recruited adults between 30 and 65. The question you asked was about adults generally, and the oldest group is not represented in anything ranked above 60.

  • Evidence gapG2

    Long-term adherence

    Nine sources discuss whether people keep using the machine. Eight of them are commentary or patient testimony. One trial measures it, and it stops at twelve months.

  • DisagreementG5

    Cardiovascular benefit

    Two well-scored sources reach opposite conclusions on whether treatment reduces cardiovascular events. Nothing in your result set reconciles them, and neither is retracted.

  • Currency gapG3

    Home testing accuracy

    The only sources on at-home diagnostic accuracy were written before the current generation of devices, and before the guidance on them was revised.

S05Actionable search suggestions

What the tool actually tells you to do next

Not autocomplete. Each suggestion is a query you could have typed yourself, the move it represents, and one sentence on why this search rather than the one you ran.

Keyword

sleep apnea

Question

Does treating it actually improve daytime function in adults?

  1. R01Split the condition
    obstructive sleep apnea

    Your results mix obstructive and central sleep apnea. They have different causes, different treatments and largely separate literatures. Naming the subtype removes the thirty-one sources that are about the other one.

  2. R02Name the population
    obstructive sleep apnea adults over 70 outcomes

    The gap map flagged that every trial you were given recruited adults between 30 and 65. If your question includes older adults, this is the query that goes looking for them.

  3. R03Name the intervention
    CPAP adherence obstructive sleep apnea

    Searching for treatment returns overviews of the options. Searching for adherence returns work with something measurable in it, which is what you need to answer whether treatment helps in practice rather than in principle.

  4. R04Name the outcome and the design
    obstructive sleep apnea daytime sleepiness randomized controlled trial

    Daytime function and cardiovascular endpoints are two separate bodies of research, and yours is the first. Adding the design term lifts primary trials above the commentary currently sitting between them.

  5. R05Use the field's own vocabulary
    obstructive sleep apnoea OR OSAHS OR "sleep-disordered breathing"

    British spelling, the older OSAHS acronym and the umbrella term sleep-disordered breathing each unlock indexed literature that the phrase you typed will never reach on its own.

Read all 9 suggestions for this search
S06From the closed beta

What people said once they had used it on their own work

Notes from beta participants, attributed by role rather than by name at their request.

  • I used to judge a source by how fast I could tell it was junk. The split score does that in about a second. Then the gap map tells me the thing I would have missed for another two days, which in my last search was that every trial I had found was about the wrong age group.

    Evidence synthesis leadPublic health agency
  • The refinement queue is the part I did not expect to care about. It is not autocomplete. It hands me a query and a sentence saying what it changes, and about a third of the time I disagree with it, which is the point, because now I know exactly what I am disagreeing with.

    Research librarianUniversity medical school
  • Our team argued about sources constantly. Now the argument happens in one place with the signals on screen and it takes ten minutes instead of a meeting. Half the time we discover we were disagreeing about the population rather than about the evidence.

    Policy analystRegulatory affairs
  • I do market research, not medicine, and I assumed this was built for academics. The independence signal alone justifies it. It flags the review sites that are selling the thing they are reviewing before I lose a morning to them.

    Independent consultantMarket and competitive research
S07Questions

The things people ask before they trust a scoring tool

Where do the sources come from?

Open web search, open academic indexes, preprint servers, guideline bodies, government and statistical publications, and news archives. There is no private walled index, and no source class is excluded by default. If a filter you set removed anything, the results header says how many and offers to put them back.

Does it decide what is true?

No, and it is built to avoid pretending otherwise. It estimates how much weight a claim can carry, which is a different question. A low trust score means a source is hard to verify, thinly evidenced, unexamined or contradicted. It does not mean the source is wrong. Plenty of correct things get published in low-scoring places. The score tells you how much checking you still owe.

What if I disagree with a score?

Override it. Any signal can be adjusted for your workspace with a note explaining why, and the override travels with the search into exports and shared links. Recording your disagreement is more useful than hiding it, particularly when somebody reviews your work later.

Is this only for medicine and science?

No. The signals are general: who is behind a source, what it is built on, whether anyone checked it, how old the claim is, who paid for it, and whether anything independent agrees. That applies to a market report, a policy consultation, a legal commentary or a trade publication as readily as to a clinical trial. Team workspaces can reweight the signals for how their field actually works.

Can I use it for a systematic review?

For scoping one and for keeping one alive, yes. The reproducible history, the diff on re-run and the retraction watch are built for exactly that. It is not a substitute for a protocol-driven database search and we would not want it presented as one. Use it before the protocol, and again afterwards to check nothing has moved.

What happens to my searches?

They stay in your workspace. They are not sold and not used to train models, and share links can be revoked at any time. Free workspaces keep searches for thirty days, Pro keeps them indefinitely, and Team workspaces set their own retention period with scheduled deletion.

Do I have to change how I search?

No. Type the keyword you were going to type anyway. The optional question line makes the scoring sharper and the gap map considerably more useful, but the tool starts from your words and shows you where they could go further rather than making you learn a syntax first.

How long does a search take?

The ranked list with both scores lands in a few seconds. The gap map takes longer, because it reads your question against the whole result set rather than each source in isolation. It appears under the results when it is ready instead of holding up the list.

Start here

Run one search on a question you already know the answer to

It is the fastest way to work out whether the scoring matches your judgement. Take a topic you know well, see what the gap map says, and decide from there. The free plan gives you twenty-five searches a month and does not ask for a card.

Questions before you sign up? The contact page reaches a person, not a queue.

Expires in

Limited time offer

We rebuilt your site for you. Claim it and we handle everything transfer, hosting, and your domain. Then update it anytime, just by asking AI.

Host for only$8 per monthBilled yearly
Claim limited offer now