In 2024, US FDA attributes 30% of recalls by all medical device manufacturers to design issues. The most common FDA 483 observations cite missing signatures, incomplete batch records, and inaccurate data. These are not isolated incidents. They point to a structural problem that manual review processes have not fully resolved.
The need
QMS assessments and records reviews are the standard assurance activities. The work is time-consuming, dependent on experienced staff, and prone to oversight when teams are under pressure.
Stay ready for regulatory changes
Quality management systems need regular assessment to remain effective, not a one-time setup. Periodic reviews follow a set schedule, while event-driven assessments respond to specific triggers such as a process deviation or a shift in the regulatory landscape. Together, they confirm that a QMS still meets current requirements and give teams the chance to plan remedial actions before small gaps become compliance failures. Internal assessments are a routine part of QMS management, and their absence often raises questions of its own.
Manage compliance risk across the product lifecycle
Records assessment checks regulatory compliance both before a product reaches the market and after it is already in use. A pre-release review evaluates records for compliance risks while a product is still in development, when adjustments cost less. A post-release review picks up issues that only appear once a product is in real-world use. Either way, teams can plan remedial actions before a gap becomes a larger problem.
How the system works

The system runs two parallel tracks. The first evaluates QMS documents against current regulatory requirements and flags gaps for remediation. The second assesses records on a pre-release and post-release basis to catch compliance risks before they escalate.
Both tracks accept documents in standard formats and evaluate them against a structured requirements matrix. Initial validation used IEC 62304, the medical device software lifecycle standard, which spans 95 clauses and 357 individual requirements. The system also covers ISO 13485, FDA 21 CFR Part 820, and the EU MDR.
The compliance engine works in three layers. The first extracts structured meaning from quality documents through semantic parsing, context-aware chunking, and text embedding. The second connects document content to specific regulatory clauses using vector search, a knowledge graph, and full-text matching. The third maps requirements to document evidence, extracts source-cited proof, and assigns a confidence score to each clause-level finding.
Outputs include a compliance report with pass, fail, or partial ratings, a gap analysis with prioritised remediations, evidence chains linking documents to specific clauses, and audit-ready exports in Excel and PDF.
Results and validation
The system achieved 97.5% accuracy on QMS document assessments reviewed by subject-matter experts, covering all 357 IEC 62304 requirements. Repeatability across multiple runs ranged from 99.8% to 100%.
Repeatability matters for a specific reason. A system that returns different results on the same documents offers limited assurance. One that converges on the same findings across repeated runs gives teams a measurable, auditable record.
Strong accuracy on a single run is necessary. Strong repeatability alongside it is what makes a process dependable.
The architecture includes human-in-the-loop validation. Experts review AI findings before the system logs or reports any result, keeping the process within the oversight framework that medical device organisations require.
Deployment considerations
Consistent results from an AI compliance system depend on how well the working context is prepared. The way the system encodes requirements and document structures directly affects output quality. Poorly constructed context produces errors that repeated runs will not correct.
Running assessments multiple times is also worth doing independently. Multiple runs let teams establish repeatability as a measured outcome rather than an assumption, and surface variance that a single run would not reveal.
Every deployment involves a trade-off among speed, accuracy, and cost. Organisations need to calibrate that balance against their regulatory environment and risk profile.
AI Agents reduce the risk of non-compliance
AI agents do not replace quality professionals. They handle the repetitive, high-volume parts of document review, reduce the risk of manual oversight, and give teams a consistent, quantifiable coverage record. For organisations still running compliance reviews manually, that combination has a direct bearing on recall exposure.
Validation conducted against Design controls (125 requirements) and IEC 62304 (95 requirements, 357 sub-requirements). The 97.5% accuracy figure reflects the accuracy of assessments reviewed by subject-matter experts and an alternate tech stack LLM-run comparison. The repeatability range of 99.8–100% is based on controlled available test data. Results in production environments may vary. US FDA data on 2024 medical device recalls and 483 observations cited as published.

