CardCite is built to get every part of a citation right, so that less of it has to be fixed by hand. Each field is read from the source itself and checked against the rest of the page, and every change to that process is tested on thousands of real pages before it is released. Below are the results of a test against three other citation tools, followed by how each field is read.
of CardCite citations were correct in every field, across the 150 web pages in the test below.

CardCite and three other citation tools were tested on 150 web pages. Websites tend to structure their pages in similar ways, so a tool that can cite one page from a site can usually cite the rest. Because of this, no website was included more than once in the test. Additionally, the 150 pages were drawn from 12 different families, such as news outlets, think tanks, government agencies, and universities, so the results reflect accuracy across the web rather than on any one kind of page.
The output of each extension was graded against the information on the page. A citation was marked incorrect if it contained false information, such as picking a publication date from a different article, or if a field was blank when it should have been filled. A blank field did not automatically count as incorrect, as some pages did not include all possible information, such as a page with no author given.
MyBib, Scribbr, and Zotero were run on those same pages, and with the same standard. Formatting, such as italics or the order of the fields, was not graded.
Each bar is the share of citations with nothing wrong in them, meaning all fields (primarily title, author, publisher, and date) that the page provided information for were filled correctly.
Install counts read from the Chrome Web Store, September 2026.
Across the same 150 graded pages.
| Field | CardCite | MyBib | Scribbr | Zotero |
|---|---|---|---|---|
| Title | 100.0% | 82.7% | 78.7% | 76.0% |
| Author | 97.3% | 72.0% | 60.0% | 56.7% |
| Publisher | 99.3% | 68.0% | 75.3% | 58.0% |
| Date | 96.7% | 69.3% | 58.7% | 59.3% |
When CardCite gets a page wrong, the fix is written as a general rule rather than a patch for that one site, so it also changes the answer on pages nobody looked at. Before a fix is kept, three tools check what it did.
None of the three tools is allowed to run on the 150 benchmark sites. The exclusion covers each whole domain, since a bug found on one page of a site tunes every page of it, and it is applied on every run automatically. The accuracy above is therefore measured on sites the detection was never developed against.
A separate program runs CardCite over sites it has never been tested on and flags anything a fixed rule can recognize as likely wrong, such as a date in the future, a title with the site name still attached, or an author that is not a person. Each flagged page is then examined to confirm the bug is real. Its largest single pass covered 34,374 pages.
503 documents, 340 web pages and 163 PDFs, frozen with the verified values each should produce. Each one is kept for a detection route, so the set tracks the routes the code can take. A change lands on a route rather than on one site, so replaying all 503 against it catches anything that used to work and stopped.
The same pages, read twice. Any change that loosens a rule is run over a frozen archive of more than 5,000 saved pages and 3,500 PDFs, none of them curated, once with the change and once without, and every page whose citation came out different is read. Where the pinned cases test the routes, this shows the change's effect across a wide sample of the web.
Below is a real differential run for one small change to how the publisher is chosen. Every saved page was cited twice, once with the change and once without, and 57 of the 3,189 came out different.
The archive
What the two runs produced
Three of the 57.
A name a page printed for itself was taken at its word. But site furniture reads like a name: a nav word, a button label, the alt text on a photo, so now and then the citation named one of those instead of the publisher.
A name like that has to survive a vote. Where two other places on the page agree on a different name, the printed one loses, and the built-in list of domains supplies the outlet's own spelling.
The old value was not the publisher at all. The field had been naming a widget, a menu word, or the alt text on a photo.
The right publisher, carrying something that is not part of its name, or given in a form the outlet does not use.
The outlet's own spelling, taken from the built-in list, in place of the shouted or over-capitalized form the page had written.
CardCite does not use AI to detect or create any part of a citation.
Every field is produced by static code, fixed rules written and tested in advance, which run the same way on every page.
The details come from the source itself, read from the page's own data and its visible text, or drawn from public catalogs of published work when the page leaves one out.
Every field is built the same way: gather every signal the page carries for it, check those signals against each other, and then determine the correct value. CardCite doesn't lean on any single part of a page, and for each field it reads the metadata in the source code alongside the text a reader actually sees. All of it has been tuned against thousands of real articles and documents.
Dates are shown in many formats: machine timestamps, written dates, or even a relative label like "3 days ago". CardCite reads all of these forms, drawing on both the page's structured data and its visible text. CardCite also allows you to choose between citing an updated date (with "Updated" added to the citation) if it is shown on the page, or to always have the publication date used.
Author lines often include text that isn't a name: staff labels, photo credits, sponsor tags. Every author candidate is screened against a list of more than 122,000 words that appear in ordinary English but not in real names, compiled from large word-frequency corpora and cross-checked against name datasets so that real but unusual surnames still pass.
Junk that fails the screen is dropped, and any real name left is kept. An author line can also carry real people who aren't the writer: when someone is credited under a production role, with a label like "Photos by", "Illustration by", or "Translated by", CardCite detects that it is a credit and keeps that name out of the author field.
Every citation names the outlet it came from. CardCite reads the publisher from the page and checks it against a built-in list of more than 200,000 domains, so an article at nytimes.com is cited as The New York Times, in the wording and capitalization the outlet itself uses.
The list also fixes what sites write in their own code, where publisher names often arrive lowercase or shortened. For a site the saved list does not cover, CardCite gathers the names the page gives itself, in its data, its masthead or logo, its copyright line and the tail of its title, and assigns weights to each one based on what signals agree and where they come from. This determines which one is chosen if any of them conflict.
A PDF is opened with the same engine Chrome uses to display one, so modern, compressed and encrypted files can all be read, with a lighter built-in parser for the damaged ones. From there the detection follows the same process and the same standards as it does on a web page, down to the same list of words that are never names and the same list of domains for naming a publisher.
A report prints its own date at the front: on the cover, or on the first pages after it. A date further in usually belongs to something the document describes, so CardCite reads the date from the front of the document and never from the body text.
CardCite is the only citation tool that marks whether a field rests on a weaker signal, and explains the values where the source or selection choice was not obvious, making it easier to see why a citation was filled in the way that it was.
Three of the four amber notes.
An amber note marks a value that rests on a weaker signal, or one carrying the sign of a common mistake. It does not mean that the value is wrong, just that it is worth checking.
Three of the seven context notes.
A quiet note explains where a value came from or why it was the one chosen, such as a date taken from the page's data because none is visible on the page, an updated date cited in place of the publication date, or where the exact day came from when the page shows only a month and year. It is there for reference, and the value beside it can usually be used as it is.
150 web pages, one page each from 150 different sites, so no site is counted twice and no site can carry the result. They span 12 families of source and 68 subject topics: wire and national news, local and regional news, government agencies, think tanks and policy institutes, law reviews, academic journals, university news offices, trade and specialist press, corporate newsrooms, health and medical news, international organizations and NGOs, and expert commentary.
The pages were gathered by a script that worked through search results and kept the first result per site that passed every check: the link returns a web page and not a file, the page is the article and not a bot wall, it carries a publication date, and its address is not a section or an index. Nothing was hand picked.
Grading did not happen inside the extensions. A program worked down the list of links, and for each one it opened the page, ran all four tools on that same loaded page, and wrote what each tool returned into a single file.
The four results were shown side by side with the tool names hidden and the panel order shuffled on every link, so a grade could not be handed to a name.
Name order, date format, capitalization, a leading "The", and URL variants that resolve to the same document are never counted as errors. Three of the four tools apply a style after extraction, and marking them for style choices would measure the style rather than the extraction. What is graded is whether the value is the one the source actually supports, which the standard sets out field by field.
The headline number counts citations where the author, date, title, and publisher are all right. A single wrong or missing field failed that citation.
The grades these figures come from are published in full, one row per page with each tool's verdict and the fields that failed, together with the standard they were graded against.