Home / Orthopedics & Spine



Artificial intelligence software is already helping clinicians detect fractures, measure alignment, and support surgical planning in orthopedics. But reaching the European market as software as a medical device takes more than strong technical performance. These tools need to show clear value in real clinical practice and meet the expectations of both the Medical Device Regulation and the developing Artificial Intelligence Act. As a CRO working in AI SaMD trials across orthopedics and robotics, we help start ups and manufacturers build credible evidence packages from the outset.
One of the most common mistakes in AI development is starting with the algorithm rather than the clinical problem. Regulators look at things differently. They want to understand who will use the software, where it will be used, what problem it is solving, and what the consequences are if it gets things wrong.
That means asking simple but important questions early. What is the clinical problem, and what benefit is the software expected to bring? A fracture detection tool might aim to reduce missed tibial plateau fractures in emergency settings. An alignment measurement tool might aim to improve consistency in surgical planning.
Who will use it, and in what setting? Software used by junior clinicians in urgent care carries a different level of risk from software used by specialist orthopaedic surgeons in a planned setting.
What happens if it is wrong? Missing a hip fracture can delay surgery and worsen outcomes. Overcalling a benign finding may lead only to additional imaging. Those differences matter because they shape classification, study design, and evidence expectations.
When intended use, user group, and clinical risk are clearly defined from the start, the evidence strategy becomes much stronger.
Why Rule 11 matters early , EU MDR classification for AI SaMD
Notified Body oversight required. Requires a compliant QMS and a clinical evidence package demonstrating benefit-risk acceptability.
Higher scrutiny. More robust clinical evidence expected. Common in surgical planning, triage, and diagnostic support tools in orthopedics.
Strictest clinical evidence requirements. Comprehensive clinical investigation likely required. PMCF architecture must be designed from day one.
Many AI orthopedic tools fall into a higher class than teams first expect. Early classification assessment prevents costly evidence restructuring.
For most AI SaMD products, classification is not something to leave until later. It directly affects the regulatory route, the level of scrutiny, and the type of evidence that will be expected.
Under MDR Rule 11, software used to support diagnostic or therapeutic decisions is often classified as Class IIa or higher. If incorrect decisions could lead to serious deterioration or surgery, the software may move into Class IIb. If the consequences could be death or irreversible deterioration, it may be Class III.
Many start ups assume software without hardware should fall into Class I. In practice, that is often not the case. In orthopedics and robotics, software used for diagnosis, triage, treatment support, or surgical planning will commonly need notified body oversight, a compliant quality management system, and a more robust clinical evidence package than teams first expect.
For AI based medical devices in Europe, the question is no longer whether MDR applies or whether the AI Act applies. For many products, both matter.
This is important because the AI Act adds expectations that sit alongside the existing device framework. In practical terms, teams now need to think more carefully about data governance, human oversight, transparency, logging, robustness, cybersecurity, and ongoing monitoring. These should not be treated as separate regulatory tasks added at the end. They need to be built into the product and evidence strategy from the beginning.
For higher risk products, notified bodies will increasingly look at how these elements are handled as part of the overall conformity assessment.
Show that a radiographic measurement correlates with osteoarthritis severity or implant size. Review literature linking fracture patterns to surgery timing.
Unit and system tests using diverse image archives. Comparing AI measurements with manual measurements. Verifying consistent output across hardware and software versions.
A reader study showing AI-assisted radiographers reduce variability in implant alignment measurement. A trial measuring how AI triage shortens referral times for suspected fractures.
Strong accuracy results are important, but they are only one part of the evidence picture. For CE marking, manufacturers need to show more than good technical performance. They need to show that the software makes clinical sense, works reliably, and adds value in real practice.
| Evidence component | Purpose | Practical examples |
|---|---|---|
| Valid clinical association | Demonstrates that the software's output is clinically meaningful for the targeted condition or decision. Evidence may come from literature, professional guidelines or exploratory studies. | Show that a radiographic measurement correlates with osteoarthritis severity or implant size. Review literature linking fracture patterns to surgery timing. |
| Technical (analytical) performance | Shows that the software processes input data accurately, reliably and precisely. Verification activities should test performance across different scanners, institutions and demographics, and document accuracy, sensitivity and specificity. | Unit and system tests using diverse image archives. Comparing AI measurements with manual measurements. Verifying consistent output across hardware and software versions. |
| Clinical performance | Demonstrates that using the software leads to improved or faster clinical decisions. Endpoints should be relevant , time to diagnosis, reduction in missed fractures, decreased inter-observer variability or impact on referral patterns. | A reader study showing that AI-assisted radiographers reduce variability in implant alignment measurement. A trial measuring how AI triage shortens referral times for suspected fractures. |
Weak endpoint thinking Insufficient for CE marking Strong endpoint thinking NB-credible evidence Impact on treatment planning Once the overall framework is clear, the next step is choosing a study design that answers the right questions. Accuracy matters, but it is rarely enough on its own. In orthopedics and robotics,

Credible AI SaMD trials do not come together through software development alone. They need input from clinical, regulatory, statistical, and technical teams from an early stage.
In orthopedics and robotics, that often means bringing together surgeons, radiologists, software teams, biostatisticians, and regulatory specialists around the same table. This early input helps define endpoints that are genuinely meaningful, strengthens risk management, and reduces the chance of costly changes later.
Patients also have a role. Their input can help teams understand which outcomes matter most, what level of trust and explanation is needed, and how the tool fits into real care pathways.
From a regulatory perspective, early discussion is just as valuable. Teams need quality systems that cover not only traditional device and software requirements, but also areas such as data governance, bias management, logging, and human oversight. Many foundations may already sit within ISO 13485 and IEC 62304, but AI based products usually need an extra layer of thinking and control.
This is often where a strong CRO partnership becomes valuable. The goal is not just to run a study. It is to help shape an evidence strategy that is clinically grounded, operationally practical, and acceptable from the start.
One of the biggest differences between AI software and traditional devices is that software does not stand still. Models may be retrained, code may be updated, bugs may be fixed, and external components may change over time. That is why manufacturers need to think about lifecycle control from the very beginning, not once the product is already on the market.
A strong change control plan should make clear how updates will be reviewed, documented, and validated. It should define which changes can be handled within the existing evidence framework and which changes may require fresh testing or additional clinical evidence.
Post market surveillance also needs to be built in early. For AI SaMD, this means more than routine vigilance. It means tracking how the software performs in real use, listening to user feedback, watching for signs of drift or bias, and feeding that information back into risk management and product oversight.
Human oversight remains central throughout. Users need to understand what the software is doing, when to trust it, and when to question it. Clear instructions, sensible safeguards, and practical escalation pathways all help make the product safer and more usable.
Cybersecurity and robustness also matter. If a product is going to support diagnosis, planning, or workflow decisions, manufacturers need confidence that it performs reliably, can cope with error conditions, and is protected against misuse or attack.
The strongest AI SaMD programs in orthopedics and robotics usually have one thing in common. They do not treat clinical evidence, regulation, and product development as separate workstreams. They bring them together early.
They start with a real clinical problem and a clear understanding of who will use the tool, where it will be used, and what difference it is expected to make. They classify the product early and build the right quality systems around it. They develop evidence in layers, showing not only technical performance but also clinical relevance and practical value. They involve clinicians, patients, and regulatory experts early enough to shape the development path, not just review it later. And they plan for what happens after launch, including monitoring, updates, and ongoing performance management.
For start ups and MedTech teams, the real challenge is not simply building an algorithm that performs well in development. It is building one that can stand up in clinical practice, regulatory review, and commercial adoption. The teams that do this well treat evidence as part of product strategy from day one.
Former Senior Clinical Reviewer at TÜV SÜD Notified Body. Expert in Class IIb/III implantable device clinical evidence, AI SaMD under EU MDR 2017/745, and Annex XIV compliance for orthopedic and spine programs.
Need help building a credible AI SaMD evidence package for CE marking in orthopedics or robotics? Eclevar MedTech designs EU MDR-compliant evidence strategies from classification to PMCF.
The people who build your evidence have sat on the other side of the table.

Whitepapers and publications produced by our teams with our notified body partners.
The services and topics connected to this page.
Send us the device, the claim and the deadline. You get a written answer within 24 hours.
Your documents are reviewed confidentially. An NDA can be signed before we receive any technical or clinical information.