Ortho / rachis

How Start Ups and MedTech Teams Can Build Credible AI SaMD Evidence for CE Marking in Orthopedics and Robotics

Home / Orthopedics & Spine

Ortho / rachisSaMD / logicielCorporate / equipe
How Start Ups and MedTech Teams Can Build Credible AI SaMD Evidence for CE Marking in Orth
WHAT THIS PAGE COVERS
1
Start with the patient, not the model
2
Get classification right early
3
MDR and the AI Act now need to be read together
4
Why accuracy alone is not enough
European CRO for medical devices, aligned with Regulation (EU) 2017/745.
Trusted by manufacturers

Leading medical device teams work with Eclevar

Terumo
Nihon Kohden
Vygon
Coloplast
Asahi Intecc

AI SaMD · Orthopedics & Robotics

CMO , Orthopedics & Spine, Eclevar MedTech

Artificial intelligence software is already helping clinicians detect fractures, measure alignment, and support surgical planning in orthopedics. But reaching the European market as software as a medical device takes more than strong technical performance. These tools need to show clear value in real clinical practice and meet the expectations of both the Medical Device Regulation and the developing Artificial Intelligence Act. As a CRO working in AI SaMD trials across orthopedics and robotics, we help start ups and manufacturers build credible evidence packages from the outset.

From product idea to credible CE evidence

  • What regulators expect from AI SaMD evidence in Europe
  • Why accuracy alone is not enough
  • How to design more credible studies early
  • What start ups can do to avoid evidence gaps

Start with the patient, not the model

Intended user, care setting, clinical workflow

What problem does it solve?

Clinical need, decision point, expected benefit

What happens if it is wrong?

Patient risk, workflow impact, regulatory consequence

One of the most common mistakes in AI development is starting with the algorithm rather than the clinical problem. Regulators look at things differently. They want to understand who will use the software, where it will be used, what problem it is solving, and what the consequences are if it gets things wrong.

That means asking simple but important questions early. What is the clinical problem, and what benefit is the software expected to bring? A fracture detection tool might aim to reduce missed tibial plateau fractures in emergency settings. An alignment measurement tool might aim to improve consistency in surgical planning.

Who will use it, and in what setting? Software used by junior clinicians in urgent care carries a different level of risk from software used by specialist orthopaedic surgeons in a planned setting.

What happens if it is wrong? Missing a hip fracture can delay surgery and worsen outcomes. Overcalling a benign finding may lead only to additional imaging. Those differences matter because they shape classification, study design, and evidence expectations.

When intended use, user group, and clinical risk are clearly defined from the start, the evidence strategy becomes much stronger.

Get classification right early

Why Rule 11 matters early , EU MDR classification for AI SaMD

Software supports diagnosis or treatment decisions

Notified Body oversight required. Requires a compliant QMS and a clinical evidence package demonstrating benefit-risk acceptability.

Wrong decision could lead to serious deterioration or surgery

Higher scrutiny. More robust clinical evidence expected. Common in surgical planning, triage, and diagnostic support tools in orthopedics.

Wrong decision could lead to death or irreversible deterioration

Strictest clinical evidence requirements. Comprehensive clinical investigation likely required. PMCF architecture must be designed from day one.

Many AI orthopedic tools fall into a higher class than teams first expect. Early classification assessment prevents costly evidence restructuring.

For most AI SaMD products, classification is not something to leave until later. It directly affects the regulatory route, the level of scrutiny, and the type of evidence that will be expected.

Under MDR Rule 11, software used to support diagnostic or therapeutic decisions is often classified as Class IIa or higher. If incorrect decisions could lead to serious deterioration or surgery, the software may move into Class IIb. If the consequences could be death or irreversible deterioration, it may be Class III.

Many start ups assume software without hardware should fall into Class I. In practice, that is often not the case. In orthopedics and robotics, software used for diagnosis, triage, treatment support, or surgical planning will commonly need notified body oversight, a compliant quality management system, and a more robust clinical evidence package than teams first expect.

MDR and the AI Act now need to be read together

For AI based medical devices in Europe, the question is no longer whether MDR applies or whether the AI Act applies. For many products, both matter.

This is important because the AI Act adds expectations that sit alongside the existing device framework. In practical terms, teams now need to think more carefully about data governance, human oversight, transparency, logging, robustness, cybersecurity, and ongoing monitoring. These should not be treated as separate regulatory tasks added at the end. They need to be built into the product and evidence strategy from the beginning.

For higher risk products, notified bodies will increasingly look at how these elements are handled as part of the overall conformity assessment.

Why accuracy alone is not enough

01 , Clinical association

Does the output make clinical sense?

Link the measurement or prediction to a real clinical decision

Show that a radiographic measurement correlates with osteoarthritis severity or implant size. Review literature linking fracture patterns to surgery timing.

02 , Technical performance

Does the software work accurately and reliably?

Reliable across datasets, users, and clinical sites

Unit and system tests using diverse image archives. Comparing AI measurements with manual measurements. Verifying consistent output across hardware and software versions.

03 , Clinical performance

Does it make a meaningful difference in practice?

Faster diagnosis or more consistent care

A reader study showing AI-assisted radiographers reduce variability in implant alignment measurement. A trial measuring how AI triage shortens referral times for suspected fractures.

Strong accuracy results are important, but they are only one part of the evidence picture. For CE marking, manufacturers need to show more than good technical performance. They need to show that the software makes clinical sense, works reliably, and adds value in real practice.

Evidence componentPurposePractical examples
Valid clinical associationDemonstrates that the software's output is clinically meaningful for the targeted condition or decision. Evidence may come from literature, professional guidelines or exploratory studies.Show that a radiographic measurement correlates with osteoarthritis severity or implant size. Review literature linking fracture patterns to surgery timing.
Technical (analytical) performanceShows that the software processes input data accurately, reliably and precisely. Verification activities should test performance across different scanners, institutions and demographics, and document accuracy, sensitivity and specificity.Unit and system tests using diverse image archives. Comparing AI measurements with manual measurements. Verifying consistent output across hardware and software versions.
Clinical performanceDemonstrates that using the software leads to improved or faster clinical decisions. Endpoints should be relevant , time to diagnosis, reduction in missed fractures, decreased inter-observer variability or impact on referral patterns.A reader study showing that AI-assisted radiographers reduce variability in implant alignment measurement. A trial measuring how AI triage shortens referral times for suspected fractures.

Choose endpoints that reflect real practice

Weak endpoint thinking Insufficient for CE marking Strong endpoint thinking NB-credible evidence Impact on treatment planning Once the overall framework is clear, the next step is choosing a study design that answers the right questions. Accuracy matters, but it is rarely enough on its own. In orthopedics and robotics,

Choose endpoints that reflect real practice

Bring experts, patients, and regulators in early

Credible AI SaMD trials do not come together through software development alone. They need input from clinical, regulatory, statistical, and technical teams from an early stage.

In orthopedics and robotics, that often means bringing together surgeons, radiologists, software teams, biostatisticians, and regulatory specialists around the same table. This early input helps define endpoints that are genuinely meaningful, strengthens risk management, and reduces the chance of costly changes later.

Patients also have a role. Their input can help teams understand which outcomes matter most, what level of trust and explanation is needed, and how the tool fits into real care pathways.

From a regulatory perspective, early discussion is just as valuable. Teams need quality systems that cover not only traditional device and software requirements, but also areas such as data governance, bias management, logging, and human oversight. Many foundations may already sit within ISO 13485 and IEC 62304, but AI based products usually need an extra layer of thinking and control.

This is often where a strong CRO partnership becomes valuable. The goal is not just to run a study. It is to help shape an evidence strategy that is clinically grounded, operationally practical, and acceptable from the start.

Plan for updates and monitoring from day one

One of the biggest differences between AI software and traditional devices is that software does not stand still. Models may be retrained, code may be updated, bugs may be fixed, and external components may change over time. That is why manufacturers need to think about lifecycle control from the very beginning, not once the product is already on the market.

A strong change control plan should make clear how updates will be reviewed, documented, and validated. It should define which changes can be handled within the existing evidence framework and which changes may require fresh testing or additional clinical evidence.

Post market surveillance also needs to be built in early. For AI SaMD, this means more than routine vigilance. It means tracking how the software performs in real use, listening to user feedback, watching for signs of drift or bias, and feeding that information back into risk management and product oversight.

Human oversight remains central throughout. Users need to understand what the software is doing, when to trust it, and when to question it. Clear instructions, sensible safeguards, and practical escalation pathways all help make the product safer and more usable.

Cybersecurity and robustness also matter. If a product is going to support diagnosis, planning, or workflow decisions, manufacturers need confidence that it performs reliably, can cope with error conditions, and is protected against misuse or attack.

What strong AI SaMD programs do differently

Clinically meaningful endpoints

The strongest AI SaMD programs in orthopedics and robotics usually have one thing in common. They do not treat clinical evidence, regulation, and product development as separate workstreams. They bring them together early.

They start with a real clinical problem and a clear understanding of who will use the tool, where it will be used, and what difference it is expected to make. They classify the product early and build the right quality systems around it. They develop evidence in layers, showing not only technical performance but also clinical relevance and practical value. They involve clinicians, patients, and regulatory experts early enough to shape the development path, not just review it later. And they plan for what happens after launch, including monitoring, updates, and ongoing performance management.

For start ups and MedTech teams, the real challenge is not simply building an algorithm that performs well in development. It is building one that can stand up in clinical practice, regulatory review, and commercial adoption. The teams that do this well treat evidence as part of product strategy from day one.

CMO , Orthopedics & SpineEclevar MedTech

Former Senior Clinical Reviewer at TÜV SÜD Notified Body. Expert in Class IIb/III implantable device clinical evidence, AI SaMD under EU MDR 2017/745, and Annex XIV compliance for orthopedic and spine programs.

Need help building a credible AI SaMD evidence package for CE marking in orthopedics or robotics? Eclevar MedTech designs EU MDR-compliant evidence strategies from classification to PMCF.

Orthopedics & Spine PMCF , EU MDR Clinical Evidence

PMCF Under EU MDR 2017/745 , Complete CRO Guide 2026

Clinical Evaluation Report (CER) Under EU MDR , Complete Guide 2026

  • MDCG 2025-6. Interplay between MDR, IVDR, and the AI Act
  • MDCG 2020-1. Guidance on Clinical Evaluation and Performance Evaluation of Medical Device Software
  • Congenius. Navigating Clinical Evaluation for SaMD
  • General Digital. MDR Rule 11 and software classification
  • Quickbird Medical. AI Act guidance for medical device manufacturers
  • Start with the patient, not the model
  • Get classification right early
  • MDR and the AI Act together
  • Why accuracy alone is not enough
  • Choose the right endpoints
  • Bring experts in early
  • Plan for updates from day one
  • What strong programs do differently
Expertise and recognition

A European team of former notified body reviewers

The people who build your evidence have sat on the other side of the table.

EUCROF Platinum Award 2026
EUCROF Platinum Award 2026xShare Open Call for Clinical Research, co-funded by the European Union
Dr Mark Da CostaDr Mark Da CostaChief Medical OfficerTÜV SÜD
Dr Nikhil KhadabadiDr Nikhil KhadabadiHead of Clinical EvidenceEU MDR
Pierre-Marie BoutanquoiPierre-Marie BoutanquoiDirector of Clinical OperationsISO 14155
Official content

Our content, signed Eclevar.

Whitepapers and publications produced by our teams with our notified body partners.

Whitepaper by BSI and Eclevar on the EU MDR
Whitepaper · BSI × Eclevar

A BSI and Eclevar whitepaper on the EU MDR.

Written with Notified Body BSI: a practical reading of the clinical evidence expectations under EU MDR 2017/745, the same evidence your file has to support.

Start the conversation

Tell us where your evidence stands today

Send us the device, the claim and the deadline. You get a written answer within 24 hours.

Your documents are reviewed confidentially. An NDA can be signed before we receive any technical or clinical information.

Reforming Clinical Evaluation of Medical Devices in Europe