About IQVIA
IQVIA provides scientific services spanning clinical trials, real world evidence, and consulting in all areas of the product lifecycle. Our Clinical Outcomes Assessments (COAs) organisation leads the industry in generating data to ensure that the patient voice is incorporated into the development and commercialisation of medication and other drug/non-drug interventions.
Role Summary
The COA Accelerator ecosystem is expanding its AI-enabled capabilities to help internal teams and external clients generate evidence-backed COA strategy recommendations. These capabilities depend on high-quality, traceable, well-governed, AI-ready data assets. The Data Engineer will play a critical role in transforming fragmented clinical, regulatory, scientific, and proprietary content into structured, searchable, and secure knowledge assets that power AI-enabled COA strategy workflows.
Responsibilities
- Design, build, and maintain data infrastructure that supports IQVIA’s AI-enabled COA strategy and COA Accelerator capabilities.
- Own the ingestion, transformation, normalisation, enrichment, indexing, versioning, and governance of public, proprietary, and client-specific data sources.
- Build ingestion pipelines for structured and unstructured sources, including PDFs, Word documents, slide decks, spreadsheets, databases, APIs, clinical trial registries, regulatory documents, scientific publications, and internal repositories.
- Transform raw source material into standardised, searchable, AI-ready formats that support evidence retrieval, source citation, recommendation generation, and expert review workflows.
- Develop repeatable processes for document parsing, OCR, text extraction, metadata enrichment, chunking, deduplication, versioning, indexing, and quality control.
- Provide technical support to the teams building and maintaining the platform’s core knowledge layer.
- Support integration of public data sources such as clinical trial registries, FDA labels, EMA EPARs, HTA records, scientific literature, FDA guidance, public qualification documents, and other relevant evidence repositories.
- Prepare data for retrieval-augmented generation workflows through high-quality chunking, embeddings, indexes, metadata filters, and source reference structures.
- Collaborate with AI engineers to improve retrieval precision, recall, relevance, and citation accuracy.
- Implement hybrid retrieval approaches combining semantic search, keyword search, structured database queries, and metadata filtering.
- Maintain traceability between AI-generated outputs and source documents.
- Implement data quality controls to identify incomplete, outdated, duplicated, poorly parsed, incorrectly tagged, or otherwise unreliable content.
- Maintain audit trails for source ingestion, transformation, updates, deletions, access rights, and downstream use.
- Work with legal, security, compliance, product, and domain stakeholders to ensure data use aligns with contractual, licensing, privacy, intellectual property, and governance requirements.