FAQ
Questions teams ask about data quality
What is the difference between data validation and data cleaning?Validation refers to the automated or predefined checks that flag potential errors at entry, such as missing values or out-of-range measurements. Cleaning is the broader process of reviewing, investigating and resolving discrepancies throughout the study, including query management, coding, external reconciliation and manual clinical review.
When should data cleaning begin?As soon as the first participant is enrolled. Continuous review distributes the workload and means issues are found while they can still be corrected, which is what shortens database lock.
Who is responsible for data cleaning?The clinical data manager coordinates it, working with investigators, clinical research associates, medical monitors, statisticians and the sponsor.
Does data cleaning modify the original clinical data?No. Corrections are made through documented queries and recorded in the audit trail. The original entries are preserved, which is what Good Clinical Practice and the ALCOA+ principles require.
How many queries is too many?When sites start answering without reading, there are too many. Query volume is not a quality metric, and a high rate often points at the edit checks rather than the data.
Can you assess a study that is already running?Yes. The assessment looks at open query age, missing critical data, external reconciliation status, coding progress and validation documentation, and produces a remediation plan with a realistic lock date.