Journals retracted more than 10,000 papers in 2023, a record at the time. The dramatic total was driven largely by more than 8,000 withdrawals from one publisher's journal portfolio after investigations of manipulated publication processes. That episode is real, but it is not a clean annual measure of how much misconduct began in 2023.
Retractions are dated when notices appear. The underlying papers may have been published across many earlier years, and a batch investigation can move thousands of accumulated cases into one calendar year. Counts also rise as total publication grows and detection improves.
The right reading holds two facts together. Coordinated fraud, paper mills, compromised peer review, and weak editorial controls have produced serious damage; at the same time, retraction is one of the mechanisms by which science identifies and marks unreliable work. A high count can reflect both a failure of prevention and an increase in correction.
What a retraction means#
Retraction tells readers that a publication should no longer be relied upon as originally presented. Under COPE guidance, reasons can include unreliable findings from major error or fabrication, falsification, plagiarism, redundant publication, unauthorized data, legal problems, unethical research, compromised peer review, or a serious undisclosed conflict.
It is more consequential than an ordinary correction. A correction repairs a limited problem while leaving the central work reliable. An expression of concern warns readers during an unresolved investigation. Removal is reserved for exceptional legal, privacy, or safety circumstances because preserving the scholarly record is important. The notice should explain the reason, identify who is retracting, and link to the article. It should remain accessible and distinguish honest error from misconduct where facts permit.
Annual counts attach to the notice date#
Imagine a publisher reviews special issues from 2018 through 2022 and publishes 2,000 notices in 2026. A chart records 2,000 retractions in 2026 even though no single 2026 publication created the batch.
This lag makes a retraction-year graph a measure of correction activity, not incidence of bad research in that year, and publication-year cohorts answer a different question but remain incomplete for recent years because future retractions have not yet occurred. So the first thing to ask of any trend is which date it uses, and whether it allows for right censoring: recent papers have had less time to be investigated.
Publication volume changes the denominator#
The number of scholarly articles has grown substantially. If a constant small fraction were retracted, the count would rise with output. A rate per 10,000 publications can help.
The denominator has to match the numerator: the same document types, fields, journals, databases, and years. If a chart counts conference papers in the output but only journal-article retractions in the numerator, the rate it shows is too low.
Rates also inherit lag. Dividing today's notices by today's publications mismatches old problematic papers with new output. Cohort-based rates are more interpretable but mature slowly.
The 2023 spike was concentrated#
Retraction Watch reported more than 10,000 notices in 2023, while noting that its manually curated database was still catching up, and public reporting attributed most of the total to mass actions involving Hindawi journals after systematic manipulation concerns.
Concentration matters. A global total can be dominated by one publisher investigation, journal cluster, or paper-mill network. It should not be described as a uniform doubling of misconduct across every field. Batch correction can expose years of weak editorial controls. It also shows why publishers need systems that detect shared patterns before hundreds of manuscripts pass through.
Paper mills industrialize manuscript production#
Paper mills sell manuscripts, data, and images. They sell authorship positions, peer-review manipulation, or combinations of these services. Their output can imitate legitimate structure while lacking trustworthy research underneath.
Indicators may include reused or altered images, implausible data patterns, and mismatched methods. They may include formulaic text, irrelevant citations, and suspicious author changes. They may include unusual email domains and coordinated reviewer accounts. No single indicator proves misconduct. COPE and STM guidance stresses multi-article investigation because patterns appear across manuscripts, journals, and publishers. Case-by-case review can miss the network and consume resources while new papers continue.
Manipulated peer review exploits trust systems#
Some journals invite authors to suggest reviewers. Fraudulent submissions can provide invented identities or controlled email accounts, producing favorable reviews. Brokers may compromise real accounts or coordinate review rings.
Weak identity verification, rapid editorial handling, and pressure to fill special issues create openings. A polished review report is not proof of an independent reviewer. Publisher defenses include institutional email checks, ORCID and affiliation verification, and reviewer-network analysis. They include editor oversight, conflict screening, and limits on author-suggested reviewers. These controls can add delay, so systems need proportional risk rules.
Special-issue growth created concentrated risk#
Guest-edited collections can bring domain expertise and organize useful communities. They can also distribute editorial authority across large temporary networks with inconsistent supervision.
Rapid expansion may outpace training, conflict review, reviewer checks, and quality assurance. Paper mills target predictable weak points and can submit many related manuscripts. The lesson is not that every special issue is suspect. Publishers should measure acceptance patterns, handling times, and reviewer reuse. They should measure citation anomalies, author clusters, and guest-editor conflicts, then audit outliers.
Better detection raises the visible count#
Image-comparison tools can find duplicated or transformed panels across papers. Text and citation analysis can identify templates and irrelevant-reference patterns. Statistical review can reveal impossible distributions or internal inconsistency.
Public post-publication critique allows readers to flag problems that traditional peer review missed. Open data and code make some claims easier to verify, and cross-publisher systems can connect patterns that no journal sees alone. Each of those makes existing problems visible for the first time, which is why an early rise in notices can mean a backlog is finally being worked through.
Open metadata improves counting#
Historically, retraction notices were inconsistently titled, linked, and indexed. Some lacked persistent identifiers or did not update the original article's metadata.
Crossref acquired the Retraction Watch database in 2023 and now makes curated retraction data available through production services and a downloadable file; the database is updated every working day and supplements publisher deposits. Improved discoverability can increase measured retractions even if publisher behavior is unchanged. It also helps reference tools warn readers and researchers build more complete analyses.
Retraction is not synonymous with fraud#
An author may discover a major coding error and request retraction. A sample can be contaminated, a key reagent misidentified, or a legal permission invalid. Those cases can make conclusions unreliable without fabrication.
Conversely, a vague notice may conceal the scope of misconduct. Databases often assign several reason categories, and one paper can involve both data and publication-process problems; analyses should use notice text and investigation findings rather than treating every retraction as an identical event. Fairness requires avoiding guilt by coauthorship when responsibility is unknown.
Low retraction counts can be bad news#
A journal with no retractions might publish flawless work, and it might also lack a reporting channel, ignore complaints, issue silent corrections, remove pages without notices, or be absent from major databases.
Institutions differ in willingness and capacity to investigate. Publishers differ in notice policy, legal risk tolerance, and metadata quality. Detection is not evenly distributed, so a low observed count cannot establish high integrity. What tells you about a journal is how it handles a credible concern.
High counts can also reveal prevention failure#
The correction explanation should not minimize harm. Thousands of unreliable papers can contaminate reviews, waste research funds, distort hiring and citation metrics, and influence clinical or policy claims.
Mass retractions show that prepublication checks and editorial governance failed at scale. Delayed notices allow more citations and derivative work, and authors whose valid papers share a journal can face reputational spillover as well. Correction is necessary after that failure, but prevention is better, and a system that reports only one of the two is telling you half the story.
Country and institution rankings are fragile#
Retraction records may list every country in a multi-author affiliation. Fractional versus full counting changes results. Institutions merge, names vary, and visiting appointments complicate attribution.
Field mix matters because publication practices, document types, and detection tools differ. A country concentrated in heavily screened biomedical fields may look worse than one concentrated in fields with sparse indexing. Publisher mix, language, collaboration, output growth, and investigation capacity also confound rankings. Rates need uncertainty intervals and transparent cleaning rules, not a simple league table.
Journal-level comparisons need maturity#
A new journal has little historical time for retractions. A journal that inherits or acquires a portfolio can appear responsible for papers accepted under earlier systems. One large batch can dominate a small denominator.
Useful measures include publication-year cohort rate, median time to notice, and reason distribution. They include share of vague notices, correction and concern rates, and proportion of notices with machine-readable links. Ask for acceptance volume and editorial structure alongside the rate. A raw total is not an answer; it tells you where to ask the next question.
Retraction lag is a safety metric#
Time from credible concern to visible notice affects downstream harm. Investigations need due process, evidence preservation, author response, and sometimes institutional findings. Speed cannot replace fairness.
Interim expressions of concern can protect readers when the issue is serious and resolution will take time. Publishers should define escalation criteria, assign responsibility, and update notices as facts develop. Long unexplained delays deserve scrutiny. A system can report many eventual retractions while allowing unreliable work to shape a field for years.
Notice quality determines what readers learn#
"This article has been retracted at the authors' request" does not tell you whether the data are unreliable, whether a permission failed, or whether a duplicated publication was removed. Vague notices hinder downstream decisions and unfairly invite speculation.
COPE recommends notices based on proven facts, with enough specificity to distinguish reasons and responsibility, and notices should be free to read and linked in both directions with the original record. A notice written that way serves safety and fairness at once: it marks the evidence as invalid without asserting more than the investigation established.
Downstream correction is the missing half#
Marking the original article does not automatically repair reviews, guidelines, or models. It does not repair patents, educational material, or AI training corpora that used it. Citation systems need status propagation, and authors need a process for reassessment.
Evidence syntheses should rerun analyses when retracted data contributed. Guidelines should review whether recommendations change. Databases and retrieval systems should retain the record but block silent evidentiary use.
The true correction time extends until affected conclusions are identified and updated, not merely until a notice DOI exists.
What stronger prevention looks like#
Publishers can verify author and reviewer identity, screen images and text, and inspect data availability. They can audit special issues, share cross-journal signals, and train editors. Institutions can improve data stewardship, supervision, authorship accountability, and protected reporting.
Funders and employers can reduce incentives that reward article count without regard to quality. Journals can adopt Registered Reports and result-neutral review where suitable. Communities can credit replication, correction, data curation, and negative results. No screening tool should issue an automatic misconduct verdict. Detection produces leads; trained humans investigate with context and due process.
How to read a retraction trend responsibly#
Begin with the unit: notice, article, conference item, preprint, or partial retraction. Identify the date convention, the database, its coverage, the duplicates, and the update type. Then match an appropriate publication denominator and allow for lag.
Break the total down by publisher, journal, field, year of original publication, and reason. Look for batches and for changes in detection policy. Treat recent cohorts as incomplete.
Then judge the correction system itself: speed, notice specificity, and metadata linkage. Judge author fairness, downstream updates, and preventive change. Until you have supplied that context, the chart you are looking at is a number, not a finding.
Sources and further reading
- Crossref documentation for the open Retraction Watch database
- Retraction Watch review of the record retraction year 2023
- COPE formal retraction guidelines
- COPE and STM research report on paper mills
- COPE supplemental guidance for multi-article paper-mill cases
- STM Integrity Hub description of cross-publisher screening
Questions and answers
Does a rising retraction count prove science is becoming less reliable?
No. It can reflect more problematic work, more publications, stronger detection, batch investigations, better metadata, or several of these at once.
Why did 2023 have so many retractions?
The record total was driven largely by mass withdrawals from Hindawi journals after systematic publication-process investigations, not a uniform global change.
Are most retractions caused by misconduct?
Reason distributions vary by database, field, and period. Major error can also require retraction, and one notice may list several causes.
Is a journal with no retractions more trustworthy?
Not necessarily. It may have excellent prevention or weak detection and correction. Transparency and response to credible concerns matter.
What is a better metric than raw count?
Use matched cohort rates plus retraction lag, reason and notice quality, metadata linkage, downstream repair, and evidence of preventive improvement.