What we measure (and what we don't)
Fluentry measures AI literacy: your ability to recognise AI outputs, evaluate their quality and reason about when and how AI is appropriate to use. Specifically, can you spot AI-generated material, do you understand how AI works at a level that lets you read the news critically and can you reason about ethical situations involving AI.
What Fluentry does not measure
- Your ability to build AI. We are not a programming test or an ML-engineering certification.
- Your ability to use a specific AI tool. We do not certify proficiency with ChatGPT, Claude, Copilot or any other product. Tools change every quarter; literacy lasts.
- Your ability to replace a human task with AI. We do not grade you on automation skill.
- Your intelligence, your employability or your future earnings potential. AI literacy is one skill among many.
If you need an assessment of one of the above, Fluentry is not the right tool. We would rather tell you that than oversell the test.
How items are made
Every Fluentry test item is co-authored. A named human educator briefs the item. Where it helps quality and speed, AI tools draft material. Two humans review every item before it ships. Where the answer depends on judgement, on our Ethics scenarios, a panel of credentialed experts sets the graded values. AI never sets a graded value.
We disclose AI's role at the item-type level. We do not hide it, and we do not overclaim it. The pipeline lives in our content repository, which is available under NDA for procurement review.
What AI does, by test
Can AI Fool You? The human side is authored by a named contributor, with no AI tool used in drafting or editing it. The AI side is generated through a rotation of five or more frontier models. Both sides are reviewed by a second named human before they merge.
AI Knowledge: the stem and the choice of concept are authored by in-house writers or commissioned subject-matter experts. AI may suggest the wrong-answer options, concept tags and explanations; the human author selects, edits and signs off. AI never writes the stem and never chooses the correct answer.
AI Ethics: stems and option phrasings are written by humans. AI may suggest variants, but the human author always writes the final. The graded value of each option is set by an expert panel of at least two practitioners. AI never sets a graded value.
What AI never does
- AI never sets a graded value on a Scenario item.
- AI never chooses the correct answer on a Knowledge item.
- AI never authors a human side on a pair test.
- AI never reviews or signs off on an item. Both reviews are human.
- AI is not used to generate imagery of real or named people, or to imitate any specific photographer's style.
- AI is not used to generate photographs of Fluentry authors, reviewers, customers or team.
Where the human photographs come from
The human side of our image pairs uses real photographs taken by working photographers. They include images by Martin Vorel, published through his free photography project LibreShot. We credit photographers here rather than on the test surface, since naming the source next to a live pair would hint at the answer.
How we score each test
Your score is the percent of items you answered correctly, with slightly more credit for harder items and slightly less for easier ones, on a 0 to 100 scale.
For our Ethics test the sentence shifts: your Ethics score is the average value of the options you picked, on a 0 to 100 scale. Higher-value options reflect the stronger ethical positions our expert panel ranked.
There is no single right answer on most Scenario items. The value scale reflects the panel's expert judgement, not an objective truth. We explain who sits on the panel further down this page.
What a Grade means
Fluentry uses a Grade scale modelled on music exam Grades: Grade 1 through Grade 8, with a Diploma tier above Grade 8.
Your Grade is computed from the per-test scores you have earned. We use your latest score on each test (not your best-of), weighted toward recent performance so a Grade you earned a year ago does not outshine the test you took last week.
The Grade thresholds are anchored against thousands of Fluentry sessions. They do not move. Grade 5 in 2027 means the same skill as Grade 5 in 2026, the same property that makes a Grade 5 violinist's certificate meaningful across decades.
What sits behind each Grade
| Grade | Represents | What it means |
|---|---|---|
| Diploma | Top 2% | The elite tier. Distinctive performance across the bank. |
| Grade 8 | Top 6% | Distinguished. |
| Grade 7 | Top 15% | Strong, with reach. |
| Grade 6 | Top 25% | Strong. |
| Grade 5 | Top 40% | Solid. |
| Grade 4 | Top 55% | Capable. |
| Grade 3 | Top 70% | Developing. |
| Grade 2 | Top 85% | Foundational. |
| Grade 1 | Below Grade 2 | Beginner. The first rung. |
These are the percentile bands at the time of the launch calibration snapshot. They are then converted to fixed score cuts and frozen; they do not roll as new takers arrive. A Grade earned in 2027 reflects the same skill as one earned in 2026.
Why anchored, not rolling
Rolling thresholds, where the cut for Grade 5 shifts each month as new takers arrive, produce grade inflation or deflation depending on which way the population shifts. Anchored thresholds produce a stable credential. We chose the credential. When we substantially change our methodology, for example by adding a new test type that affects how Grades aggregate, we explicitly relock and notify every user whose Grade changes. This is the same discipline music exam boards have used for decades.
The Diploma tier
Diploma is the top 2% of all Fluentry Grades. It is not Grade 9; it is a separate tier with its own credential treatment. We chose 2% because it is rare enough to feel meaningful and not so rare it disappears in any given launch year.
Future direction: when our Academy ships, the Diploma tier may evolve to require course completion in addition to the score threshold. We will announce that change in advance and update this page on the same day it ships.
When a new test ships
A user who took only one Fluentry test has a Grade based on that test. When a new test ships, their Grade does not change overnight; they keep what they earned. They climb (or hold, or slip) when they take more tests. We never force a retake.
How we compare you
Comparing your score against other takers in your country only works once enough of them have taken it. If your country is still small, we tell you. Your share card always shows a percentile, we will never leave you without a number, but we disclose where we drew it from: your country, the wider European field or (while worldwide numbers are still building) the global field.
The share image you generate freezes at the moment you share. If your country gains more takers next month and your local percentile shifts slightly, the picture you posted to LinkedIn does not retroactively change.
Who you are compared against
| When | Source | What you see |
|---|---|---|
| Your country has enough takers | Your country | Better than X% of takers in {country} |
| Your country is still small but Europe is large | European field | Better than X% of European takers |
| Both are still small | European field (banded) | Top X% of European takers so far |
| Worldwide numbers are still building | Global (banded) | You're among the first to take this; the comparison is still forming |
An honest caveat
When we combine European scores across countries, the result is an approximation. The item banks differ: French pair-test items are not translations of English ones; they are independently authored to a shared difficulty rubric. The combined European figure is therefore an aggregate of comparably-difficult-but-not-identical tests. As each country's own numbers grow past their first few hundred takers, we switch back to your country's own field, which is the cleanest comparison.
Item health
Fluentry items are continuously evaluated against the people who take them. Each week, every active item is rated on accuracy and engagement. Items that get too easy, too hard or that show a sudden jump in accuracy (a sign the answer has leaked) are flagged. Items that consistently fail to discriminate retire from the active bank.
Aggregated item-health figures for the active bank are available under NDA for procurement review.
When an item stops working, it gets too easy because AI improves, it shows up on a public forum or it just does not reveal anything useful any more, it retires. We replace it. The active item bank turns over continuously. We log every retirement and report aggregated counts under NDA for procurement review.
We aggregate this data per test and per test type, never per item (item identity is intellectual property we do not surface).
Who builds the bank
Fluentry's item bank is built by a team that combines in-house writing, commissioned subject-matter expertise and a network of educators, ethicists and policy practitioners. Every item is reviewed by a second human before it ships. For Ethics scenarios, the graded values are set by an expert panel of at least two practitioners.
We do not publish individual names on the public site, most contributors are commissioned per item and that is their normal preference. We do disclose the standard we hire to. For procurement review, we provide the full author and reviewer list under NDA.
The standard we hire to
- Authors: in-house writers with working subject knowledge, or commissioned experts with published or practising experience in the test domain. Pair-test human sides are written by native speakers of the test's language.
- Reviewers: a second human, independent of the author, a native speaker of the test's language, with enough working knowledge of the domain to assess accuracy and fit.
- Expert panel (Ethics): at least two panellists per item with practising or published experience in AI ethics, AI policy, philosophy or adjacent fields. The median of the panel sets the graded value.
We do not name individuals publicly at v1 because most contributors prefer it, because procurement reviewers get more value from verifying the roster under NDA and because we do not publish personal data we do not need to.
Change log
When the way Fluentry measures changes, we record it here. Three kinds of change get logged. Item changes: items retire continuously as the bank evolves; we log the quarterly batch refreshes, not every individual retirement. Methodology changes: when we add a new test type, change how Grades aggregate or revise the percentile thresholds, this page gets a versioned entry, and anyone whose Grade is affected gets an email before the change is reflected. Page changes: when this page itself is meaningfully revised, the entry below records what changed and when.
| Date | Scope | Summary |
|---|---|---|
| 3 June 2026 | Page | Editorial revisions for clarity across several sections. No change to methodology, scoring or Grade thresholds. |
| v1 launch | Page | Methodology page first published with three live tests (Can AI Fool You?, AI Knowledge, AI Ethics) in en-GB. |
EU AI Act position
The EU AI Act came into force in February 2025. Two parts of it are directly relevant to Fluentry.
Article 4, the AI literacy obligation. Organisations using AI systems must ensure their staff have sufficient AI literacy. Fluentry exists, in part, to make that obligation measurable. Our B2B offering documents organisational AI literacy in a form suitable for procurement and audit.
High-risk classification (Annex III). The Act lists categories of AI systems that count as high-risk and trigger heavier compliance obligations. Fluentry is not high-risk under Annex III. We are an AI literacy assessment used by individuals on themselves and by organisations on their own staff for compliance purposes. We are not used for hiring decisions, credit scoring, biometric identification, education-grading in the regulated school or university sense or any other Annex III category.
To be clear: a Fluentry Grade is not a regulated educational certificate. It is a credential we issue against our own methodology, modelled on music exam Grades but not part of any national education system, awarding body or accredited qualification framework. It belongs in the same conceptual category as a language-app streak or a typing-speed badge: a real signal of a real skill, issued by the platform that measures it, not a state-recognised academic award.
Where a B2B customer chooses to use Fluentry in a way that does engage Annex III obligations, for example by treating our assessment as a binding criterion in a hiring decision, that use is the customer's responsibility under the Act, not Fluentry's. Our standard B2B contracts allocate those obligations through the data processing agreement.
An internal compliance memo with the full Annex III analysis is available to procurement teams under NDA.
ACADEMY METHODOLOGY
How the Academy works
The Academy teaches to the gaps our tests measure, against recognised public frameworks, with every course expert-reviewed and dated. This page explains how it is built and why you can trust it.
What AI literacy means here
We use one model across the Tests and the Academy, so what you are measured on is what you are taught.
AI literacy is more than knowing what a model is. We assess and teach across four pillars: Knowledge (what AI is and how it works), Use (getting useful, safe results from it), Critical (judging output, spotting failure and bias) and Ethics (the responsible and lawful choices around its use).
The Academy is the remediation layer on top of the Tests. The test shows where you are weak, the Academy teaches to that gap and a re-take proves the gain on the same scale.
The frameworks we build against
Every course and role path maps to recognised public frameworks. This is the procurement-grade answer to the question a buyer or regulator asks first: what standard is this against. The frameworks are public, so the mapping is free credibility rather than a private claim.
| Our coverage | Mapped frameworks |
|---|---|
| All-staff foundations | EU Commission Article 4 guidance; DigComp 3.0 (AI dimension); UNESCO AI competency frameworks |
| Use and prompting | DigComp 3.0; EC-OECD AILit framework |
| Critical evaluation and risk | EC-OECD AILit framework; UNESCO AI competency frameworks |
| Ethics, oversight and the law | EU Commission Article 4 guidance; UNESCO Recommendation on the Ethics of AI |
| Manager and oversight role path | EU Commission Article 4 guidance; EC-OECD AILit framework |
Frameworks evolve. We review the mapping each quarter and date the page so you can see when it was last checked.
How a course is made
Courses are drafted with AI from real data on what learners get wrong, then a named expert reviews and signs off, then the course is translated and checked, then it is reviewed and dated every quarter. We disclose AI co-authoring openly, the same stance we take on the Tests, because transparency is a trust asset with EU buyers rather than a liability.
- AI drafts from the gap data: the item-level signal of what people actually get wrong.
- A named subject expert reviews, corrects and signs off before anything goes live.
- Translation and a native-language check across our launch locales.
- A dated quarterly review, with a per-course changelog.
How we measure learning
Every module ends in an assessment built on the same engine as the Tests. We record competence, not just completion, which is the evidence advantage over tick-box training. Completing a course is not the credential; demonstrating the competence is.
Learning feeds the credential. After working a path you re-take the relevant test, your score recomputes your overall Grade and the Grade is the thing that climbs. The credential stays earned, not farmed.
Our EU AI Act position
The Academy is a training product. Its AI features (the practice sandbox and the explain-why tutor) are educational tools, not consequential automated decisions about a person. Our preliminary assessment is that the Academy is not a high-risk AI system, and AI use is clearly disclosed throughout.
This is our preliminary position. The definitive statement publishes here once external counsel has confirmed it.
The evidence and the credential
For an organisation, the Academy produces an EU AI Act Article 4 evidence record: a dated, per-person, per-course log of training and competence, with the leaver records retained, that you can show a regulator. It records competence, not just attendance.
The credential itself is the Grade, not a wallet of course badges. It is one living credential that you come back to improve and that gently decays if you stop, so it reflects current competence. That is the difference between measured fluency and a certificate of attendance.
LAST REVIEWED 6 June 2026
