Skip to main content
Rubric v1.06 product types

How we score, and why you can check it

Every score on this site is a weighted average of published criteria. The weights live in code, so a review cannot re-weight itself. We show you the components, so if you value something differently you can re-derive your own total.

The scale

Nought to ten. Individual criteria are scored in half points; the weighted total is published to one decimal place. Each band has a written anchor, so the same number means the same thing across every review and every product type.

  • 9.0 - 10ExceptionalBest in class. It replaced the thing we were already using.
  • 8.0 - 8.9ExcellentRecommended without reservation for the audience it targets.
  • 7.0 - 7.9GoodRecommended, with caveats we name explicitly.
  • 6.0 - 6.9FairIt works. A better option exists and we tell you which.
  • 5.0 - 5.9MediocreOnly if a specific constraint forces your hand.
  • 3.0 - 4.9PoorWe do not recommend it.
  • 0.0 - 2.9BrokenActively fails at its job, or harms the person using it.

Colour on this site means score and nothing else: 8.0 and above, 6.0 to 7.9, and below 6.0. Never category, never product type.

The rubrics

A phone app and a keyboard cannot be judged on the same axes. Each product type has its own weighted rubric, and every set sums to exactly 100 — the build fails otherwise.

Android & iOS apps

Weighted toward whether the app does its one job, and toward what it asks of your phone and your data in exchange.

  • 30%Core job performanceDoes the primary task work, correctly, quickly, every time9: The task is one tap from the launcher and never fails. 4: The primary task needs workarounds or fails intermittently.
  • 20%Usability & first runTime to first value, onboarding friction, discoverability9: Useful within a minute, no account wall, no tour needed. 4: Signup required before anything works; core actions hidden.
  • 15%Reliability & performanceCrashes, cold-start time, jank, battery and thermal cost9: No crashes across the test window; sub-second cold start. 4: Crashes in normal use, or measurably drains the battery.
  • 15%Privacy & permissionsPermissions asked vs needed, trackers, export and delete9: Only permissions it uses; no trackers; data exportable. 4: Permissions unrelated to function; no way to get data out.
  • 10%Value & monetisationPrice against alternatives, paywall aggression, dark patterns9: Fair price, honest paywall, nothing manipulative. 4: Core function paywalled after the fact, or dark patterns.
  • 5%AccessibilityScreen reader, dynamic type, contrast, offline behaviour9: Fully operable with a screen reader at large type sizes. 4: Primary interface invisible to assistive technology.
  • 5%Support & maintenanceUpdate cadence, changelog quality, developer responsiveness9: Regular updates, human-written changelogs, replies to bugs. 4: Abandoned, or updates that only add monetisation.

Websites & SaaS

Heavier on data portability and commercial terms, because the cost of a bad web tool is mostly the cost of leaving it.

  • 28%Core job performanceDoes the product do the work it claims to do9: Handles the real workload without workarounds. 4: Breaks down at realistic data volumes.
  • 20%Usability & information architectureCan a new user find and finish the critical path9: Critical path is obvious and completes without docs. 4: Core task buried; navigation fights the user.
  • 16%Privacy, security & portabilitySSO, 2FA, export, subprocessors, DPA availability9: Full export in an open format; 2FA standard; DPA published. 4: No export, no 2FA, subprocessors undisclosed.
  • 14%Performance & reliabilityMeasured LCP, INP, CLS; status page and incident history9: Fast on measurement; incidents disclosed and resolved. 4: Sluggish under normal use; outages go unacknowledged.
  • 12%Pricing & commercial termsTransparency, seat traps, lock-in, how hard it is to cancel9: Public pricing, no seat traps, cancel in the product. 4: Contact-sales-only pricing and a retention maze to leave.
  • 5%AccessibilityWCAG spot-check plus a keyboard-only run of the critical path9: Critical path completes keyboard-only with visible focus. 4: Keyboard traps, or focus never visible.
  • 5%Support & documentationResponse times, docs quality, community9: Accurate docs; support answers within a working day. 4: Docs stale or absent; support unreachable.

Windows, macOS & Linux

Similar to mobile, but resource use and telemetry matter more when software runs all day on a work machine.

  • 28%Core job performanceThe primary task, at realistic project size9: Handles large real projects without stalling. 4: Degrades badly once the project is non-trivial.
  • 20%Usability & learnabilityDiscoverability, keyboard support, defaults9: Sensible defaults; everything reachable by keyboard. 4: Essential features undiscoverable without a tutorial.
  • 16%Performance & resource useMemory, CPU at idle, startup, behaviour under load9: Modest idle footprint; starts fast; stays responsive. 4: Heavy at idle, or makes the machine unusable under load.
  • 12%Privacy & telemetryWhat it phones home, and whether you can turn it off9: Telemetry disclosed and genuinely disableable. 4: Undisclosed collection with no opt-out.
  • 12%Value & licensingPerpetual vs subscription, offline activation, seat rules9: Fair licence, works offline, reasonable seat terms. 4: Requires constant connectivity to stay licensed.
  • 6%AccessibilityPlatform accessibility APIs, contrast, scaling9: Respects platform accessibility settings throughout. 4: Custom UI invisible to platform assistive tech.
  • 6%Support & update cadenceRelease rhythm, regression history9: Predictable releases; regressions fixed quickly. 4: Long gaps, and updates that break existing work.

Physical products

Led by measurement against the manufacturer's own claims, then by whether the thing survives being used.

  • 25%Performance vs claimsMeasured against the spec sheet, with method and run count9: Meets or beats every published figure under our method. 4: Materially under-performs its own spec sheet.
  • 20%Build quality & durabilityMaterials, tolerances, wear after the test window9: No measurable wear or loosening across the window. 4: Visible degradation within weeks of normal use.
  • 18%Everyday usability & ergonomicsHow it feels after hour twenty, not hour one9: Comfortable through long sessions; controls fall to hand. 4: Causes strain, or controls are unreachable in use.
  • 12%Battery, power & thermalsMeasured runtime against claimed, under stated load9: Meets claimed runtime; stays cool under sustained load. 4: Well under claimed runtime, or throttles quickly.
  • 10%Companion softwareThe app or driver, which is usually the weakest part9: Optional, stable, and does not require an account. 4: Required, unstable, and account-gated.
  • 10%Value for moneyAgainst direct competitors at the same street price9: Clearly better than anything at the price. 4: Beaten by cheaper alternatives on every axis.
  • 5%Repairability, support & warrantyParts availability, teardown difficulty, warranty terms9: Parts sold openly; sensible warranty honoured. 4: Glued shut, no parts, warranty hard to claim.

AI tools, assistants & APIs

Scored on a fixed prompt suite with a published pass rate, because a demo proves nothing and a vibe proves less.

  • 30%Task quality & accuracyFixed prompt suite, published pass rate, same suite every time9: High pass rate on the published suite, reproducibly. 4: Fails a third or more of the suite.
  • 15%Reliability & consistencySame prompt across N runs; we report the variance9: Near-identical answers across runs. 4: Materially different answers to the same prompt.
  • 15%Safety, refusals & hallucinationWhat it invents, and how confidently it invents it9: Says it does not know rather than inventing. 4: Fabricates citations or facts with full confidence.
  • 15%Privacy & data-use termsTraining on your data, retention, deletion, residency9: No training on customer data by default; deletion honoured. 4: Trains on your inputs with no opt-out.
  • 10%Latency & throughputMeasured p50 and p95 under stated conditions9: Fast and stable at p95, not just p50. 4: Unusable tail latency under normal load.
  • 10%Cost & valueReal cost at realistic volume, not list price9: Predictable cost that holds at real volume. 4: Cost balloons unpredictably with normal use.
  • 5%Integration & developer experienceSDKs, docs, error messages, escape hatches9: Clear errors, honest docs, easy to leave. 4: Opaque failures and no way out.

Subscriptions & services

For things that are neither software nor hardware: hosting, courses, managed services. Delivery and terms carry the weight.

  • 30%Core deliveryDid the service do what was promised, on time9: Delivered as promised, every time, without chasing. 4: Missed commitments with no proactive communication.
  • 18%Usability of the experienceSignup, account management, everyday interaction9: Everything self-serve and obvious. 4: Routine changes require a support ticket.
  • 15%Reliability & SLAUptime or delivery record against the stated commitment9: Meets or beats the published SLA, with evidence. 4: Misses its own SLA and does not credit it.
  • 15%Terms & transparencyContract clarity, price changes, exit terms9: Plain terms, notified price changes, clean exit. 4: Auto-renewal traps and punitive exit terms.
  • 15%Value for moneyAgainst alternatives at the same commitment level9: Clearly worth the price against alternatives. 4: Costs more for less than the obvious alternative.
  • 7%SupportReachability and competence when something goes wrong9: Reachable humans who resolve the issue. 4: Scripted responses that never resolve anything.

What a review must carry before it can publish

These are not guidelines. They are enforced by the build — a review missing any of them fails to deploy. That is deliberate: a rule you can skip under deadline is not a rule.

5+

Evidence items

At least 3 screenshots and 1 measurement, each stamped with its capture date and device.

10+

Elapsed days

A product has to be lived with, not opened once. Hours alone do not qualify.

3–8 h

Hands-on minimum

Varies by product type. Hardware needs the most; a mobile app the least.

2+

Named drawbacks

A review with no cons fails validation. Nothing is good at everything.

3+

Named strengths

Stated specifically enough to disagree with.

1

Exact version

A build string, never "latest". Software changes; the score is a snapshot.

The commitments

What we will not do

  • Score a product SofNerds built. Ever.
  • Accept payment for coverage, or for a better score.
  • Send a draft or a score to a vendor before publication.
  • Put a sponsored or gifted product in a ranked guide.
  • Publish a score without publishing its component criteria.
  • Quietly edit a verdict. Corrections are dated and logged.

Rubric versioning

Weights change. Old scores do not silently move.

Every review is pinned to the rubric version it was scored under and is always recomputed against that version. Changing a weight cannot rewrite published history. A re-score happens at a scheduled re-test, and the review says which version produced its number.

A category or tag becomes browsable once 3 products have been reviewed in it. Below that it renders for readers but stays out of search, so nobody lands on a thin page and nobody hits a dead end.