Data Sources
All data on Voxsanity is sourced from publicly available government registries. This page lists each source, its licence status, whether commercial use is permitted, and how we handle attribution.
Confirmed sources
| Source | Data type | Licence | Commercial use | Attribution |
|---|---|---|---|---|
| ClinicalTrials.gov clinicaltrials.gov/api/v2 |
Clinical trial registrations | US Federal public domain | Yes | Source linked on every trial page |
| openFDA api.fda.gov |
Drug approvals, labelling data | CC0 1.0 Universal | Yes | Not required (CC0 waives all rights) |
| NIH Reporter api.reporter.nih.gov |
Research funding grants | US Federal public domain | Yes | Not required |
| OpenAlex api.openalex.org |
Academic papers and research volume. Not currently fetched and no paper figure is published on Voxsanity | CC0 | Yes | Not required (CC0) |
| PatentsView search.patentsview.org |
Patent filings (pipeline signal). Not currently fetched and no patent figure is published on Voxsanity | US Federal public domain | Yes | Not required |
| PBS (Pharmaceutical Benefits Scheme) api.pbs.gov.au |
Australian drug subsidy listings | © Commonwealth of Australia, CC BY 3.0 AU. Use and redistribution permitted with attribution; the data must not be modified | Yes | Required on every page and export that shows PBS data |
| RxNorm / RxNav rxnav.nlm.nih.gov |
Drug name normalisation (brand to generic), mechanism of action | US Federal public domain | Yes | Credited as good practice (not legally required) |
| ISRCTN registry isrctn.com |
Clinical trial registrations (UK-based international registry) | CC0 for registry metadata; CC BY 4.0 for contribution content | Yes | Required — shown on every ISRCTN trial page, with a link to the original record |
| Australian Bureau of Statistics data.api.abs.gov.au |
Estimated Resident Population by state; SEIFA 2021 socio-economic indexes by postal area | CC BY 4.0 | Yes | Required — shown on /insights/access-equity/, the only page that uses it |
| OECD (GBARD) sdmx.oecd.org |
Government budget allocations for health R&D, for international comparison | CC BY 4.0 | Yes | Required — shown on /insights/, the only page that uses it |
Notes on specific sources
ClinicalTrials.gov
Clinical trial data is sourced from the ClinicalTrials.gov v2 API, operated by the US National Library of Medicine. Eligibility criteria are returned in CommonMark Markdown format by the API. This data is in the public domain as a work of the United States Federal Government.
The trial catalogue is checked nightly, but individual trial records are refreshed on a rolling basis rather than all at once, so a record may be up to several weeks old. Each trial page and trial card shows the date its own record was last synced. Where a record has lagged, its status, sites and enrolment can differ from the current registry state in either direction, so check the registry link on the trial page before acting on them. Trial status can change at any time. Always verify current status directly with the trial site before acting on this information.
openFDA
Drug approval and labelling data is sourced from the openFDA API, operated by the US Food and Drug Administration. openFDA data is released under the Creative Commons CC0 1.0 Universal licence, which means all rights are waived and no attribution is legally required. We link to FDA source records regardless.
Note: GMDN device data accessed via openFDA is not covered by the CC0 licence and is not used by Voxsanity.
PBS (Pharmaceutical Benefits Scheme)
PBS data is © Commonwealth of Australia and is published under a Creative Commons Attribution 3.0 Australia licence. Attribution is required, not optional: every page and export on Voxsanity that shows PBS-derived information carries the PBS attribution line alongside it, and the licence also requires that the data is not modified from its original wording where it is displayed verbatim. If you reuse PBS-derived data you obtain from Voxsanity, the same conditions travel with it — see our terms.
ISRCTN registry
ISRCTN is a UK-based international clinical trial registry. Registry metadata published from 2019 onward is released under CC0; contribution content is CC BY 4.0. Both permit commercial use, and CC BY 4.0 requires attribution, so every ISRCTN trial page names the registry as the source and links to the original record.
Australian Bureau of Statistics
Estimated Resident Population figures (dataflow ERP_Q) and SEIFA 2021 socio-economic indexes by postal area are used under CC BY 4.0, which permits commercial use with attribution. The attribution and the licence link appear directly beneath the sections that use them on /insights/access-equity/. Census and ERP population counts are different collections taken at different times and are never mixed in the same calculation.
OECD (GBARD)
Government Budget Allocations for R&D, socio-economic objective 07 (health), are used under CC BY 4.0 for the international funding comparison on /insights/, with attribution shown alongside. One honest caveat: the OECD's canonical terms page returned an HTTP 403 error when we tried to read it directly, so our CC BY 4.0 reading rests on the licence statements published on the OECD's own data pages rather than on that terms page.
PatentsView
PatentsView is not used, and we do not fetch it. Patent data is deliberately not pursued unless a real patient-facing reason for it comes up, so no patent figure is published anywhere on Voxsanity and nothing on this site rests on this source. That is a product decision rather than a technical limitation: nothing is broken and no key or migration is being waited on.
RxNorm / RxNav
Drug name data, including brand-to-generic name resolution used by the medicine search and mechanism-of-action and drug-class labels on medicine pages, comes from RxNorm via the RxNav API, produced by the US National Library of Medicine (NLM). RxNorm is in the public domain. Drug name data: RxNorm, National Library of Medicine.
How we process data
Raw data from each source is stored in our database after each nightly sync. Plain English descriptions of eligibility criteria are generated using AI language models and stored separately from the original source text. The original source text is always preserved and linked. When a plain English interpretation is not yet available, the original text is displayed.
AI-generated plain English summaries are reviewed against source data on a sample basis. If you notice a discrepancy, please contact us.
Update frequency
Trial data: updated nightly. Drug approval data: updated nightly. Pipeline and funding data (NIH Reporter): updated nightly. PBS listing data: updated monthly. The timestamp on each data record shows when it was last pulled from its source.
PBS listing data is updated monthly, not nightly, because the PBS publishes a new Schedule once a month.
Last completed sync
| Source | Schedule | Last completed without errors |
|---|---|---|
| Trial data | Updated nightly | 18 August 2026 |
| Drug approval data | Updated nightly | 4 August 2026 |
| Pipeline and funding data (NIH Reporter) | Updated nightly | 18 August 2026 |
| PBS listing data | Updated monthly | 3 August 2026 |
Each date is the most recent run of that sync that finished with no errors. A run that finished only partly is not counted, so a date here can be older than the newest data on the site, never newer. These are the four sources whose update schedule is stated above; other sources we use are loaded on their own schedules.
Data quality and errors
Voxsanity presents data as it appears in source registries. Errors or inconsistencies in source data (for example, incomplete eligibility criteria or missing dates) are shown as received. If you identify a data quality issue, please let us know and we will review it.