Methodology
Every company on TrustRating gets a star rating from 0 to 5. The rating blends an AI panel of leading large language models with human reviews, weighted by recency and verified-purchase status.
How a rating is built
Two halves of equal size. A machine panel that reads everything consistently, and the people who actually dealt with the company.
When both halves exist, neither can outvote the other. A company adored by its customers and doubted by the panel lands in the middle — and the profile shows you both sides rather than hiding the tension inside an average. When a listing has no published reviews yet, the profile shows the AI panel's assessment on its own, clearly marked as preliminary until the first verified review arrives.
A rating is only worth something if the method behind it survives being read closely.
Most rating systems ask you to trust a number. We would rather you trust the process that produces it, which means the process has to be written down in enough detail that you could argue with it. This page is that write-up: what we collect, how we weigh it, when it changes, and what deliberately has no effect on it at all.
The first rule is that no single voice decides a company's rating. A panel of large language models is queried independently, and once customers who have actually dealt with the company start reviewing it, their voice is given exactly as much weight as that panel. Neither side can drown out the other, and neither is treated as the ground truth. Until the first review arrives, the profile carries the panel's assessment alone, labelled as preliminary.
The second rule is that every input stays inspectable. We store each model's individual answer rather than a blended summary, so you can see where the panel disagreed with itself, and reviews stay attached to accounts rather than to anonymous strings. The third rule is that commercial relationships never touch the arithmetic. A company can pay us for tooling; it cannot pay us for a number.
Four stages, in order. Nothing skips a stage, and every stage is repeatable — which is why any rating on the platform can be rebuilt from its inputs.
We assemble what is publicly knowable about a company: its own disclosures, its regulatory footprint where one exists, the category it competes in, and the reviews already published on TrustRating. That material is normalised into a single structured brief so that every company is presented to the panel the same way.
The brief goes to every enabled model with an identical prompt and an identical rubric. Models never see each other's answers, so agreement between them is a signal rather than an echo, and each response is stored in full alongside the model that produced it.
Category scores are averaged across models to form the AI half. Published human reviews are averaged on the same 0 to 5 star scale to form the other half. When both halves exist they are combined in equal measure into the rating shown on the profile; while one half is still missing, the profile shows the other half on its own and marks the rating as preliminary.
We query each of the following models with a structured prompt asking them to independently score the company across five categories: trust, regulation, trading/product, support, and reputation. Each category is scored on the 0 to 5 star scale. The overall AI score is the unweighted mean of all category scores across all enabled models.
| Model | Focus |
|---|---|
| GPT-5 | OpenAI |
| Claude | Anthropic |
| Gemini 2.5 | google · 1,048,576 tok ctx |
| Grok 4.3 | x-ai · 1,000,000 tok ctx |
| DeepSeek V3 | DeepSeek |
Models and reviewers answer the same questions, in the same order, on the same scale. That is what makes a rating comparable between a bank and a courier.
The bands below are how we describe a rating in words. They are labels for readers, not thresholds that change any calculation.
Once a listing has published reviews, they count for 50% of the displayed rating (the other 50% comes from the AI panel). Star ratings from 1 to 5 are averaged directly. Only published reviews count; pending, hidden, and flagged reviews do not contribute. Reviews from invitations carrying a valid order reference get a “Verified purchase” badge.
Every published review carries the same weight in the average, and each one shows its date so you can judge recency yourself. Reviews that arrive through an invitation carrying a valid order reference are marked as verified purchases, which makes them easier for you to weigh and considerably harder for a competitor to fake.
Every one of these has been asked for at least once. The answer does not change.
Every newly-submitted review runs through a battery of heuristic checks: duplicate text, burst posting, same-IP clusters, reward offers in the body, brand-new accounts, and shared user-agent fingerprints. Above a confidence threshold we automatically promote the review to the moderation queue. We never auto-delete reviews — a human always has the final call.
TrustGuard is the screening layer between the two. It flags; people decide.
Writing a review requires a free account with a confirmed email address. Anonymous submissions and unconfirmed addresses never reach the next stage.
A battery of heuristics runs immediately: duplicate or near-duplicate text, bursts of reviews arriving together, clusters sharing an address or a device fingerprint, language offering or requesting rewards, and accounts created moments before posting.
Above a confidence threshold the review goes to the moderation queue instead of straight onto the profile. Below it, the review publishes and the signal stays on file. Nothing is ever removed automatically.
A moderator keeps, hides, or removes the review against our published guidelines. The company can respond publicly either way — but it can never edit or delete what a customer wrote.
Human reviews recompute the score on publication, edit, or removal. AI panel scans run on demand and on a rolling schedule.
In practice, a rating you read today reflects every review published up to this moment, plus a panel assessment refreshed on a schedule rather than on every page load. That is a deliberate choice: re-asking every model the same question on every visit would cost a great deal and change the answer very rarely.
TrustRating is independent. We do not sell score adjustments, we do not remove accurate reviews, and no commercial relationship changes the arithmetic described on this page.
If something here is unclear, or looks wrong to you, tell us. The method is meant to be argued with.
Because they read a great deal of public material consistently, and consistency is what makes comparison possible. A model applies the same rubric to the thousandth company as to the first, which no human panel can do at this scale. What models cannot know is what it felt like to be a customer — which is precisely what the human half of the rating is for.
Nothing entirely, which is why we keep and show each model's individual answer instead of only the average. Where the panel disagrees, you see the disagreement. Where it agrees, you can decide for yourself whether that agreement is evidence or a shared blind spot — and the human half of the rating is there to contradict it.
No. Paid plans unlock tooling — replies, review invitations, widgets, analytics — and nothing else. There is no price for a better number, no priority queue for scores, and no plan tier that changes how a review is weighed.
Every company profile shows the panel's answers and the reviews behind its rating — including the ones that disagree with each other.
The result is written to the company profile with its contributing evidence attached. Any review that is published, edited, or removed triggers a recomputation, and panel scans re-run on a rolling schedule, so a rating tracks reality instead of freezing at the moment it was first calculated.
Publication, edit, or removal recomputes the human half straight away, so the number on the profile and the reviews underneath it never disagree.
Rarely, and only to fix a demonstrable error — evidence attached to the wrong company, or a scan that ran against a stale profile. Corrections are recorded internally and exist to repair faulty inputs, never to settle a disagreement about what a company deserves.
Because the panel half does not depend on review volume. A company with two reviews still has a panel assessment, and the rating reflects it. A company with no reviews at all shows the panel's assessment on its own, marked “Preliminary — AI panel only” until the first verified review arrives. The review count sits next to the rating everywhere it appears, so you can always judge how much of the number rests on lived experience.
The human half changes the moment a review is published, edited, or removed. The panel half changes when a scan re-runs, on demand or on schedule. Ratings therefore drift continuously rather than being reset in periodic batches.