CardCite

Accuracy

CardCite is built to get every part of a citation right, so that less of it has to be fixed by hand. Each field is read from the source itself and checked against the rest of the page, and every change to that process is tested on thousands of real pages before it is released. Below are the results of a test against three other citation tools, followed by how each field is read.

92.7%

of CardCite citations were correct in every field, across the 150 web pages in the test below.

fiercehealthcare.com/providers/connecticut-doctors-now-sue-patients-most-over-medical-bills-surpassing-hospitals
CardCite's outputper field
The top of a Fierce Healthcare article: the site logo, the headline, the byline Noam N. Levey, KFF Health News, Katy Golvala, CT Mirror, Jenna Carlesso, CT Mirror, and the date Apr 20, 2026.
Publisherfrom the page data
Fierce Healthcare
Title3 of 4 sources agree
"In Connecticut, doctors now sue patients most over medical bills, surpassing hospitals"
Author3 names kept
Noam N. Levey, Katy Golvala and Jenna Carlesso

KFF Health News and CT Mirror were recognized as organization names rather than people.

Dateshown date matches the data
April 20, 2026
An article with CardCite's output for each field. MyBib, Scribbr and Zotero each got at least one of these four fields wrong on it.

How the score was reached

CardCite and three other citation tools were tested on 150 web pages. Websites tend to structure their pages in similar ways, so a tool that can cite one page from a site can usually cite the rest. Because of this, no website was included more than once in the test. Additionally, the 150 pages were drawn from 12 different families, such as news outlets, think tanks, government agencies, and universities, so the results reflect accuracy across the web rather than on any one kind of page.

The output of each extension was graded against the information on the page. A citation was marked incorrect if it contained false information, such as picking a publication date from a different article, or if a field was blank when it should have been filled. A blank field did not automatically count as incorrect, as some pages did not include all possible information, such as a page with no author given.

MyBib, Scribbr, and Zotero were run on those same pages, and with the same standard. Formatting, such as italics or the order of the fields, was not graded.

Read the full method

How the four tools scored

Each bar is the share of citations with nothing wrong in them, meaning all fields (primarily title, author, publisher, and date) that the page provided information for were filled correctly.

0%50%100%
CardCite
This extension
11 failed: 6 with a wrong value, 5 with a blank
92.7%
139 of 150
MyBib
1,000,000 Chrome Web Store users
102 failed, all with a wrong value
32.0%
48 of 150
Scribbr
1,000,000 Chrome Web Store users
115 failed: 91 with a wrong value, 24 with a blank
23.3%
35 of 150
Zotero
8,000,000 Chrome Web Store users
122 failed: 73 with a wrong value, 49 with a blank
18.7%
28 of 150

Install counts read from the Chrome Web Store, September 2026.

How often each field was correct per tool

Across the same 150 graded pages.

Field CardCite MyBib Scribbr Zotero
Title 100.0% 82.7% 78.7% 76.0%
Author 97.3% 72.0% 60.0% 56.7%
Publisher 99.3% 68.0% 75.3% 58.0%
Date 96.7% 69.3% 58.7% 59.3%
See all grades

How every change to CardCite is tested

When CardCite gets a page wrong, the fix is written as a general rule rather than a patch for that one site, so it also changes the answer on pages nobody looked at. Before a fix is kept, three tools check what it did.

None of the three tools is allowed to run on the 150 benchmark sites. The exclusion covers each whole domain, since a bug found on one page of a site tunes every page of it, and it is applied on every run automatically. The accuracy above is therefore measured on sites the detection was never developed against.

1. The scanner

A separate program runs CardCite over sites it has never been tested on and flags anything a fixed rule can recognize as likely wrong, such as a date in the future, a title with the site name still attached, or an author that is not a person. Each flagged page is then examined to confirm the bug is real. Its largest single pass covered 34,374 pages.

2. Pinned test cases

503 documents, 340 web pages and 163 PDFs, frozen with the verified values each should produce. Each one is kept for a detection route, so the set tracks the routes the code can take. A change lands on a route rather than on one site, so replaying all 503 against it catches anything that used to work and stopped.

3. The differential

The same pages, read twice. Any change that loosens a rule is run over a frozen archive of more than 5,000 saved pages and 3,500 PDFs, none of them curated, once with the change and once without, and every page whose citation came out different is read. Where the pinned cases test the routes, this shows the change's effect across a wide sample of the web.

One change, tested on 3,189 saved pages

Below is a real differential run for one small change to how the publisher is chosen. Every saved page was cited twice, once with the change and once without, and 57 of the 3,189 came out different.

An example run over 3,189 saved pages, each read twice
read 0 of 3,189 came out different 0 of 57

The archive

not read yet read, citation unchanged came out different

What the two runs produced

Page Without the change With the change
bigwaterut.gov Spinner: White decorative Big Water Town
rockrivertimes.com THE ROCK RIVER TIMES Rock River Times
macfound.org John D. and Catherine T. MacArthur Foundation MacArthur Foundation

Three of the 57.

Each cell in the grid is one of the 3,189 saved pages, in the order they were read. The 57 pages whose citation changed with the new code stay lit, and the list beside the grid shows the before and after for ten of them.
What this change was
The rule before

A name a page printed for itself was taken at its word. But site furniture reads like a name: a nav word, a button label, the alt text on a photo, so now and then the citation named one of those instead of the publisher.

The rule after

A name like that has to survive a vote. Where two other places on the page agree on a different name, the printed one loses, and the built-in list of domains supplies the outlet's own spelling.

What the 57 turned out to be
13
A different name

The old value was not the publisher at all. The field had been naming a widget, a menu word, or the alt text on a photo.

  • "Spinner: White decorative"Lansing Township
  • "Kelly Nguyen on blurred background"MRC Laboratory of Molecular Biology
  • "ESI"European Stability Initiative
6
Trimmed or expanded

The right publisher, carrying something that is not part of its name, or given in a form the outlet does not use.

  • "John D. and Catherine T. MacArthur Foundation"MacArthur Foundation
  • "THE ROCK RIVER TIMES"Rock River Times
  • "Nigerian Journal Of Water, Sanitation And Hygiene For Development (NWASHDEV)"Nigerian Journal of Water, Sanitation and Hygiene for Development
38
Capitalization only

The outlet's own spelling, taken from the built-in list, in place of the shouted or over-capitalized form the page had written.

  • "BELLS UNIVERSITY OF TECHNOLOGY"Bells University of Technology
  • "American Association For Applied Linguistics"American Association for Applied Linguistics
  • "WESTVIEW NEWS"WestView News

How each field is determined

CardCite does not use AI to detect or create any part of a citation.

Every field is produced by static code, fixed rules written and tested in advance, which run the same way on every page.

The details come from the source itself, read from the page's own data and its visible text, or drawn from public catalogs of published work when the page leaves one out.

Every field is built the same way: gather every signal the page carries for it, check those signals against each other, and then determine the correct value. CardCite doesn't lean on any single part of a page, and for each field it reads the metadata in the source code alongside the text a reader actually sees. All of it has been tuned against thousands of real articles and documents.

Dates

Dates are shown in many formats: machine timestamps, written dates, or even a relative label like "3 days ago". CardCite reads all of these forms, drawing on both the page's structured data and its visible text. CardCite also allows you to choose between citing an updated date (with "Updated" added to the citation) if it is shown on the page, or to always have the publication date used.

The page shows
Published Feb 27, 2020 | Updated March 3, 2021
The citation gets
Updated March 3, 2021

Authors

Author lines often include text that isn't a name: staff labels, photo credits, sponsor tags. Every author candidate is screened against a list of more than 122,000 words that appear in ordinary English but not in real names, compiled from large word-frequency corpora and cross-checked against name datasets so that real but unusual surnames still pass.

Junk that fails the screen is dropped, and any real name left is kept. An author line can also carry real people who aren't the writer: when someone is credited under a production role, with a label like "Photos by", "Illustration by", or "Translated by", CardCite detects that it is a credit and keeps that name out of the author field.

The author line readsThe citation gets
By Michael Ross→Michael Ross
Michael Ross and Getty Images→Michael Ross

Publishers

Every citation names the outlet it came from. CardCite reads the publisher from the page and checks it against a built-in list of more than 200,000 domains, so an article at nytimes.com is cited as The New York Times, in the wording and capitalization the outlet itself uses.

The list also fixes what sites write in their own code, where publisher names often arrive lowercase or shortened. For a site the saved list does not cover, CardCite gathers the names the page gives itself, in its data, its masthead or logo, its copyright line and the tail of its title, and assigns weights to each one based on what signals agree and where they come from. This determines which one is chosen if any of them conflict.

The page's addressThe citation gets
nytimes.com→The New York Times
dw.com→Deutsche Welle
csis.org→Center for Strategic and International Studies

PDFs

A PDF is opened with the same engine Chrome uses to display one, so modern, compressed and encrypted files can all be read, with a lighter built-in parser for the damaged ones. From there the detection follows the same process and the same standards as it does on a web page, down to the same list of words that are never names and the same list of domains for naming a publisher.

A report prints its own date at the front: on the cover, or on the first pages after it. A date further in usually belongs to something the document describes, so CardCite reads the date from the front of the document and never from the body text.

The PDF saysThe citation gets
"Published October 2025" on the cover→October 2025
"In March 2019, Congress passed a law..." on page 14→not used as the date

CardCite gives the context behind a value

CardCite is the only citation tool that marks whether a field rests on a weaker signal, and explains the values where the source or selection choice was not obvious, making it easier to see why a citation was filled in the way that it was.

Amber note
This publisher has no capital letters. Double-check its capitalization.
Date detected from page text. Please verify accuracy.
Author detected from page layout. Please verify against the page.

Three of the four amber notes.

An amber note marks a value that rests on a weaker signal, or one carrying the sign of a common mistake. It does not mean that the value is wrong, just that it is worth checking.

Quiet note
No publication date is shown on this page. This is the website's official publication date, from its data.
An updated date is shown on this page, so it was used instead of the publication date.
No author byline is shown on this page. This name comes from the website's data.

Three of the seven context notes.

A quiet note explains where a value came from or why it was the one chosen, such as a date taken from the page's data because none is visible on the page, an updated date cited in place of the publication date, or where the exact day came from when the page shows only a month and year. It is there for reference, and the value beside it can usually be used as it is.

How the test was run

1

How the sources were chosen

150 web pages, one page each from 150 different sites, so no site is counted twice and no site can carry the result. They span 12 families of source and 68 subject topics: wire and national news, local and regional news, government agencies, think tanks and policy institutes, law reviews, academic journals, university news offices, trade and specialist press, corporate newsrooms, health and medical news, international organizations and NGOs, and expert commentary.

The pages were gathered by a script that worked through search results and kept the first result per site that passed every check: the link returns a web page and not a file, the page is the article and not a bot wall, it carries a publication date, and its address is not a section or an index. Nothing was hand picked.

2

Every citation was graded blind

Grading did not happen inside the extensions. A program worked down the list of links, and for each one it opened the page, ran all four tools on that same loaded page, and wrote what each tool returned into a single file.

The four results were shown side by side with the tool names hidden and the panel order shuffled on every link, so a grade could not be handed to a name.

3

The value was graded, never the formatting

Name order, date format, capitalization, a leading "The", and URL variants that resolve to the same document are never counted as errors. Three of the four tools apply a style after extraction, and marking them for style choices would measure the style rather than the extraction. What is graded is whether the value is the one the source actually supports, which the standard sets out field by field.

4

A citation is correct only when all of it is correct

The headline number counts citations where the author, date, title, and publisher are all right. A single wrong or missing field failed that citation.

The grades these figures come from are published in full, one row per page with each tool's verdict and the fields that failed, together with the standard they were graded against.

Back to the CardCite home page