“Define the population and endpoints before collecting data, set the rules for harmonization and provenance at the design stage, and keep quality control independent.”
When There’s No Registry: Building Rare Disease Evidence That Supports Regulatory Scrutiny
Key Takeaways
- Fit-for-purpose rare disease RWE begins by defining intended regulatory use, which drives cohort criteria, endpoints, follow-up, and traceability, preventing decision-inadequate data collection.
- Lack of registries necessitates building multinational datasets with prespecified index dates and aligned endpoint definitions, while harmonizing heterogeneous documentation across notes, labs, imaging, and EHR systems.
In rare disease drug development where no registry or natural history dataset exists, real-world evidence quality depends on treating evidence engineering as a design-stage decision, defining intended regulatory use and standardizing endpoints, harmonization, and data provenance upfront rather than reconciling gaps after collection.
In rare disease drug development, the hardest external-control problem often starts before any analysis begins: there may be no registry, no mature natural history dataset and no obvious evidence source to evaluate. As real-world evidence (RWE) moves earlier in the drug development process and is used to support higher-stakes decisions, sponsors need to think less about whether real-world data belongs in clinical development and more about how to build evidence that is fit for purpose from the start.
FDA guidance has helped make that shift more concrete by recognizing the role RWE can play in supporting regulatory decision-making. But in rare diseases, that recognition creates a harder practical question: what do you do when the evidence source you need does not yet exist?
Start with the question, not the data
The strategic point is not that rare disease evidence requires more data—it's that the evidence must be engineered with its eventual regulatory use in mind. In practice, that means treating cohort definition, endpoint selection, harmonization, provenance and quality control as part of the study design, not as downstream data-management tasks.
The first design decision is therefore not which data source to use, but what the evidence must prove. An external control arm, a natural history study, a regulatory submission, and an exploratory analysis all place different demands on the dataset. Defining the intended use upfront determines who should be included, which data elements need to be captured, how much follow-up is required, and what level of traceability will be needed later.Without that clarity, teams can spend months collecting data that is clinically interesting but not strong enough for the decision it is meant to support.
Recent industry research makes a similar point, describing approaches that identify data gaps early rather than at analysis stage.
When there is no registry, design choices become evidence choices
In rare diseases, that intended-use question often leads to a different starting point than it would in more common conditions. There may be no registry to draw from, no mature natural history dataset, and no single healthcare system with enough relevant patients. The focus therefore shifts from selecting an existing evidence source to building a fit-for-purpose dataset across countries, sites, and care settings.
With no registry to base the study on, the case definition and inclusion criteria have to be set by the study team and applied consistently across countries where diagnosis may be done differently. The external control also has to be genuinely comparable to the trial arm with similar endpoints and definitions, aligned timepoints and index dates, and agreed before any data is collected, not reconciled once the gaps surface. Because these studies are usually international, data harmonization becomes a central challenge from day one. The goal is not simply to bring data together from different sources; it is to understand and manage the differences in how clinical information is captured, so it can be analyzed in a comparable way.
That is rarely straightforward. Information for a single patient may be spread across clinical notes, pathology reports, laboratory records, imaging reports, and structured EHR systems, requiring information from multiple sources to be brought together before a complete patient journey can be understood. Building a patient-level dataset is as much about interpreting these sources consistently as collecting them, and data quality, completeness and linkage are among the hardest problems in this work. Clear design upfront is essential, but a rare disease study is also a learning process so as data comes in, there is a clear understanding of variation across sites and countries. The important thing is to incorporate those learnings in a controlled and transparent way, rather than making fundamental changes late in the process.
Confidence doesn’t come from numbers
This is where rare disease evidence can challenge a common assumption about RWE. In rare diseases, scale may be impossible, so the quality question shifts: can the dataset explain who the patients are, why they are comparable to the trial population, are they representative, what information may be missing and how much uncertainty remains? Generating credible evidence becomes less about scale and more about the quality and transparency of the data.
In rare diseases, confidence rarely comes from scale. It comes from whether the cohort is precisely defined, the relevance of the data elements, completeness of follow-up, and understanding of the limitations. Representativeness matters more than it first appears, and not only in terms of demographics. It also includes disease severity, treatment history, diagnostic delays, and the referral patterns that concentrate certain patients in specialist centers. For some rare diseases, a meaningful proportion of the available population may already have taken part in clinical trials, which makes it harder to find external-control patients who reflect standard care rather than prior exposure to investigational treatment.
Missing data deserves the same scrutiny because the pattern of what is missing can be just as important as the volume. Information may be absent because of how care is delivered, because sicker patients are documented more thoroughly, or because untreated patients are less visible in the record. In small populations, there is limited ability to adjust those differences statistically, and changes in how patients are treated over time can make comparisons even harder. A robust study design does not remove these limitations, but it helps teams identify them early, assess their impact, and explain them transparently.
Building for submission from the start
Where evidence is required to support a regulatory decision, submission-grade data must be a design choice, not a final thought.FDA guidance emphasizes the use of supported data standards and clear documentation of data transformations. In practice, this often means mapping and standardizing real-world data into submission-ready structures, while maintaining appropriate documentation and traceability. However, real-world data is collected in routine clinical care, not to a predefined study schedule. Assessments may happen at different timepoints, variables may be captured differently across sites, and some information expected in a trial may not exist at all. Turning that into submission-ready data is therefore not a simple mapping exercise; it requires early decisions about what can be standardized, what must be documented, and what uncertainty will remain.
Data provenance deserves the same early attention because it is the evidence trail regulators use to judge whether the dataset can be trusted. They need to understand where the data came from, how patients were identified, how outcomes were defined and how quality was maintained. That audit trail has to be built as the data is assembled; it cannot easily be reconstructed later. In an interventional study a site can often be queried while the context is still fresh. In retrospective observational work, records may be incomplete, clinicians may have moved on, and some questions simply cannot be resolved after the fact.
These studies carry real ethical obligations too, especially because rare disease datasets can be inherently more identifiable. Data minimization—collecting only what the study genuinely needs—is both an ethical and a regulatory expectation.Rare disease studies bring additional privacy considerations. Small patient populations can increase the risk of re-identification, particularly when data is drawn from multiple sources and contains detailed clinical information. As a result, careful consideration needs to be given to anonymization, data governance, and how results are reported.
What good looks like in practice
These principles can sound procedural, but in rare diseases they directly determine whether the evidence will be usable. The practical question is how these design choices show when a team must build the dataset itself.
In one program built to support a therapy for a rare autoimmune hematologic condition, which affects roughly one in 8,000 people and had no approved treatment or registry, generating usable evidence meant creating a patient-level dataset from scratch. Data was collected on more than 300 patients across eight countries, including the UK, US, and Japan, with independent statistical quality control and traceability to submission standards. This ensured that when no mature evidence source existed, the resulting evidence was well designed, assembled and documented. That evidence went on to form part of a regulatory submission, and the therapy has since been approved by the FDA.
The scale of data was not the point. What made the evidence usable is that each feature answered a design-stage decision. The geographic spread across eight countries made harmonization essential from the start rather than something to reconcile later. Traceability supported the regulatory requirement for data provenance. Independent quality control allowed the evidence to be scrutinized without the analysis effectively checking its own work.
For teams designing these studies, the principle is easy to state and harder to achieve: start from the evidentiary standard and work backwards. Define the population and endpoints before collecting data, set the rules for harmonization and provenance at the design stage, and keep quality control independent. Investing this effort upfront rarely slows a study down; in fact, it reduces rework and helps teams answer regulatory questions with confidence because the data-generating process is well understood and documented. The aim is not to eliminate uncertainty but to generate the most reliable evidence the situation allows, and to be transparent about its strengths and its limits.
As regulators grow more comfortable with external controls and real-world data studies, the real differentiator will not be whether a team can access rare disease data. It will be whether the data has been built into evidence regulators can trust.
Susanna Lövdahl, PhD, vice president of data operations, BC Platforms
Sources
- Applied Clinical Trials, Studna A 2026. Real-World Evidence in Clinical Trials: Where the Industry Stands and Where It's Headed. Applied Clinical Trials.
https://www.appliedclinicaltrialsonline.com/view/real-world-evidence-clinical-trials-industry?ekey=YW1hbmRhQGF3Y29tbXMuY28%3D - Applied Clinical Trials, 2026. Do H, Zhang V, Lamberti MJ, Morgan C. The Use of Real-World Data and Evidence in Clinical Trials.
https://www.appliedclinicaltrialsonline.com/view/use-real-world-data-evidence-clinical-trials - Food and Drug Administration (FDA). (2023). Considerations for the Use of Real-World Data and Real-World Evidence to Support Regulatory Decision-Making for Drug and Biological Products (Guidance for Industry). FDA.
https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-real-world-data-and-real-world-evidence-support-regulatory-decision-making-drug . - Food and Drug Administration (FDA) 2023: Data Standards for Drug and Biological Product Submissions Containing Real-World Data, December 2023.
www.fda.gov/regulatory-information/search-fda-guidance-documents/data-standards-drug-and-biological-product-submissions-containing-real-world-data
Related to this article









