Evidence explainer

Digital health and AI

WHO's Six Principles for Artificial Intelligence in Health

WHO's six principles are not a badge or checklist. They are a connected governance framework for autonomy, safety, transparency, accountability, equity, and responsive, sustainable AI across a system's life.

Fully reviewed by Jasaman (Jasmin) Tojjar, MD, PhD

On this page
  1. 1. Protect human autonomy
  2. 2. Promote human well-being, human safety, and the public interest
  3. 3. Ensure transparency, explainability, and intelligibility
  4. 4. Foster responsibility and accountability
  5. 5. Ensure inclusiveness and equity
  6. 6. Promote AI that is responsive and sustainable
  7. Why the principles cannot be separated
  8. Turn the principles into a procurement matrix
  9. Measure the workflow, not just the model
  10. Relationship to law and technical frameworks
  11. References

Ethical principles are easy to place in a slide deck and hard to keep alive in a working health system. A vendor can promise fairness, a hospital can promise human oversight, and a policy can promise transparency while patients still encounter an inaccessible interface, an unsafe recommendation, or no meaningful way to challenge a decision.

The World Health Organization's Ethics and governance of artificial intelligence for health names six connected principles. They cover autonomy; well-being, safety, and the public interest; transparency, explainability, and intelligibility; responsibility and accountability; inclusiveness and equity; and responsiveness and sustainability.

These principles are not six labels to check once. They are questions you apply across problem selection, data collection, and development. You apply them across validation, procurement, and deployment. You apply them across monitoring, incident response, and retirement. Their usefulness comes from translating each one into observable evidence and a person with authority to act.

1. Protect human autonomy#

WHO's first principle says humans should remain in control of health systems and medical decisions. Privacy and confidentiality should be protected, and valid informed consent should be supported by appropriate legal frameworks.

Autonomy is more than placing a clinician near an algorithm. It asks whether a person understands the role of the system, whether a professional can reject its output, and whether patients have a meaningful route to ask questions or choose an alternative when one should be available.

An AI tool can narrow choice before anyone sees a recommendation. A scheduling model may decide who receives an appointment. A triage system may determine which message reaches a clinician. A documentation tool may frame the clinical story in language that influences later care. Human control must cover these upstream effects, not only the final click.

Consent also needs precision. General consent to treatment is not automatically consent for every secondary data use, model-training arrangement, recording, or commercial transfer. Requirements vary by jurisdiction and setting, but a team should map each data flow. It should identify the lawful basis, notice, and consent rules. It should identify retention and deletion rules.

Operational controls include role-based permissions, a non-AI pathway where appropriate, clear notices, a visible correction route, protection against retaliation for refusal, and periodic testing of whether your users can actually override the system. An override button that is hidden, slow, or punished by workflow metrics does not create meaningful autonomy.

2. Promote human well-being, human safety, and the public interest#

The second principle requires AI designers to meet safety, accuracy, and efficacy requirements for well-defined uses. Quality control and quality improvement should continue during real use.

The phrase “well-defined use” is decisive. A model that summarizes stable outpatient visits has not thereby proved it can triage chest pain. A system validated on adult chest radiographs has not proved performance in children. A general chatbot has not proved that its medication guidance is complete.

Your safety assessment should identify who can be harmed, how, and with what severity. This includes wrong answers, omissions, and delay. It includes unnecessary testing, privacy loss, and cybersecurity compromise. It includes deskilling, workload transfer, and the opportunity cost of funding the tool instead of another service.

Average accuracy is rarely enough. You need clinically meaningful endpoints, severe-error analysis, and subgroup performance. You need calibration, robustness to incomplete data, and testing under realistic workload. The reference standard itself should be examined because historical decisions may encode error or inequity.

The public-interest component widens the view beyond one institution. A tool might improve throughput for a well-resourced hospital while diverting scarce computing, data, or clinical labor from higher-priority needs, and procurement should ask whether the problem warrants AI, whether a simpler intervention works, and who receives the benefit.

3. Ensure transparency, explainability, and intelligibility#

WHO describes transparency as publishing or documenting sufficient information before design or deployment so that meaningful public consultation and debate are possible. Explainability concerns reasons or factors behind an output. Intelligibility asks whether the relevant audience can understand what is communicated.

These concepts overlap but are not identical. A highly technical model card can be transparent to an engineer but unintelligible to a patient. A simple explanation can feel understandable while concealing weak evidence. An interpretable feature list can still rest on a biased target variable.

Different audiences need different information. Regulators may need technical files and validation evidence. Procurement teams need intended use, versions, data practices, known limits, and change control. Clinicians need workflow-specific performance and response guidance. Patients need plain language about the tool's role, material limits, data use, and how to seek correction.

Generative systems add a citation problem. A model may invent a source or attach a real source that does not support the claim. Systems that present references should test link validity and claim-level support. Source visibility helps, but it does not absolve you of validating the output.

Transparency also covers uncertainty and absence. A system should communicate when input quality is poor, when the case lies outside intended use, and when no reliable answer is available; designing an honest refusal can be more important than designing another confident response.

4. Foster responsibility and accountability#

AI performs tasks, but people and institutions remain responsible for the conditions under which it is used, and WHO calls for accountability and mechanisms for questioning and redress when individuals or groups are harmed by algorithmic decisions.

Responsibility should be allocated before launch. Name an executive owner, clinical owner, technical owner, privacy and security contacts, and operational incident lead. Define who can approve updates, who can suspend the system, and who communicates with affected patients.

A vague statement that “the clinician is responsible” is inadequate when the organization selected the vendor, set the workflow, limited the time available for review, and controls whether the model can be bypassed. Accountability must match actual authority.

Contracts are part of the control system. They should address data rights, model and subcontractor changes, security duties, incident notification, audit evidence, performance commitments, version history, records access, service termination, and continuity when the vendor leaves the market.

Redress requires an accessible complaint path, timely review, and preservation of relevant logs. It requires correction of records, communication of findings, and remediation of systemic problems. A complaint should not be treated as an isolated customer-service event when it may reveal a repeated safety or equity failure.

5. Ensure inclusiveness and equity#

WHO calls for AI that supports equitable use and access irrespective of age, sex, or gender. That holds irrespective of income, race, or ethnicity. It holds irrespective of sexual orientation, ability, and other protected or relevant characteristics. This principle concerns the entire system, not only a fairness metric.

Representation in development data matters, but presence alone is not enough. Labels may reflect unequal prior care. Data quality may differ by site. Small subgroup counts can produce unstable estimates. A model can perform similarly in a benchmark while its interface remains unusable with assistive technology or its service remains unaffordable.

Inclusive design brings affected patients, health workers, disability advocates, language communities, and civil society into decisions early enough to change them. Participation should influence whether the use is pursued, which outcomes matter, which harms are unacceptable, and how fallback care works.

Evaluation should report relevant subgroup performance with uncertainty, inspect intersections where feasible, and examine allocation effects, and if an algorithm sends more resources to people with greater prior spending, it may reproduce unequal access even if it never receives race as an explicit input.

Equity controls include accessibility testing, language quality review, and alternative channels. They include representative local validation, monitoring for differential errors and service outcomes, and a funded response plan. Collecting demographic data without a purpose, protection, or remediation path can create its own harm.

6. Promote AI that is responsive and sustainable#

The sixth principle requires continuing assessment of whether a system responds to actual expectations and needs. It also includes environmental impact, energy efficiency, workforce disruption, and training.

Responsiveness means a system can be corrected, restricted, or retired when the context changes. Clinical guidance evolves. Disease prevalence shifts. A vendor updates a model. A connected data field changes meaning. A new patient group begins using the service. Validation at launch cannot cover every future state.

Change control should identify which modifications require review and revalidation. Version numbers, prompts, and retrieval sources all can affect performance. So can thresholds, user interfaces, and workflow rules. A model update described as an improvement may introduce a new failure mode.

Sustainability includes computing energy, water use, hardware, and electronic waste, but it is broader than a carbon estimate. A project that cannot be maintained, audited, staffed, or funded safely is not sustainable. Vendor dependency and loss of institutional knowledge are resilience risks.

Workforce effects deserve direct assessment. Automation can remove repetitive work, but it can also create correction labor, surveillance pressure, skill loss, and new duties without time or training; a responsive deployment measures these effects and changes staffing or scope accordingly.

Why the principles cannot be separated#

A system can be transparent about being unsafe. It can be accurate while denying patients a meaningful choice. It can protect privacy while allocating care inequitably. It can perform well at launch and degrade after an undocumented update.

That is why a single score is a poor representation of ethical acceptability. Tradeoffs need explicit reasoning. More detailed logs may improve accountability but increase privacy risk. Collecting demographic variables may enable equity auditing but requires data protection. A model's environmental cost may be acceptable for a rare lifesaving task and unjustified for cosmetic convenience.

Your governance record should state the tradeoff, affected groups, and evidence. It should state decision authority, safeguards, and review date. Saying “balanced by the committee” is less useful than documenting what was balanced and how future evidence could change the decision.

Turn the principles into a procurement matrix#

For autonomy, ask how users are informed, how data are used, which controls exist, and whether an alternative path is available. Test the override rather than accepting a screenshot.

For safety and public interest, request intended-use evidence, severe-failure analysis, and external and local validation. Request security testing and proof that the proposed problem cannot be addressed more safely or simply.

For transparency, require versioned documentation, data provenance at an appropriate level, and performance limitations. Require uncertainty behavior, model-change notices, and audience-specific explanations.

For accountability, identify owners, incident timelines, and audit rights. Identify redress, record correction, indemnity where appropriate, and the power to suspend service. Confirm that responsibility follows control.

For inclusion and equity, request accessibility evidence, language performance, and subgroup results. Request community participation methods, local evaluation, and a plan for disparities after launch.

For responsiveness and sustainability, request update governance, revalidation triggers, and continuity plans. Request workforce assessment, environmental information, exit terms, and a safe retirement process.

Measure the workflow, not just the model#

An AI result reaches people through an interface and an organization. Evaluation should simulate real cases with representative users, interruptions, incomplete records, time pressure, and downstream handoffs.

Useful measures can include serious error rate, correction burden, and time to resolution. They can include override quality, subgroup differences, and accessibility defects. They can include complaints, delayed actions, security incidents, energy use, and staff workload. Each measure needs a threshold and response.

A pilot should have a bounded population and predefined stop criteria. Expansion to a new use, specialty, language, or patient group is a new claim. It should not occur simply because the existing tool is popular.

Periodic review should ask whether benefits remain real, whether harms have appeared, whether a simpler option now exists, and whether continued use still serves the public interest. Retirement is a normal governance outcome, not necessarily a project failure.

Relationship to law and technical frameworks#

WHO's principles are normative guidance. They do not approve a product or certify a developer. They do not replace laws on medical devices, privacy, or civil rights. They do not replace laws on consumer protection, cybersecurity, professional practice, or procurement.

The NIST AI Risk Management Framework offers a complementary process organized around govern, map, measure, and manage. The IMDRF good machine learning practice principles address AI-enabled medical-device development across the product life cycle. WHO's large multimodal model guidance applies the ethical foundation to generative systems.

Map every relevant source to a single control register of your own. Duplicated principles can share evidence; conflicting or jurisdiction-specific requirements need documented resolution. A policy that only lists framework names does not show compliance or ethical performance.

References#

  1. WHO ethics and governance of artificial intelligence for health
  2. WHO announcement of six guiding principles
  3. WHO executive summary on ethics and governance of AI for health
  4. WHO guidance on large multimodal models
  5. WHO regulatory considerations on AI for health
  6. NIST Artificial Intelligence Risk Management Framework 1.0
  7. IMDRF good machine learning practice guiding principles

Questions and answers

What are WHO's six principles for AI in health?

They are protecting human autonomy; promoting human well-being, safety, and the public interest; ensuring transparency, explainability, and intelligibility; fostering responsibility and accountability; ensuring inclusiveness and equity; and promoting AI that is responsive and sustainable.

Are the six principles a certification standard?

No. They are normative guidance. They do not provide product approval, a legal safe harbor, a numerical score, or proof that a particular system is safe for a particular use.

Can an AI system satisfy one principle while failing another?

Yes, and that is why the principles must be assessed together. Strong documentation does not cure unsafe performance, and high average accuracy does not cure inaccessible design or absent redress.

What does protecting autonomy require in practice?

It requires meaningful human control, suitable notice and consent, privacy, usable correction and override paths, and an alternative where appropriate. The details depend on the use and applicable law.

How should an organization use the principles during procurement?

Turn each principle into evidence requests, contract duties, workflow tests, accountable owners, monitoring measures, thresholds, incident procedures, and conditions for limiting or retiring the tool.