Live Research Data Platform for Pharma | AgentixLake
← All case studies
CASE STUDY · TOP-10 GLOBAL PHARMA

New research, ready for NLP researchers within a minute.

For a top-10 global pharma company, our founder helped build an event-driven AWS data platform. It ingests research publications from sources such as Springer and PubMed as they are published, and serves them to NLP researchers working on new medicines.

New publications usable in under a minute1B+ Iceberg records

UPDATED OCTOBER 2026

The challenge

Research publications arrived live from publishers and databases such as Springer and PubMed, in many document formats and at changing volumes. The platform had to keep the raw evidence, normalize metadata, coordinate enrichment and make every new paper usable quickly, while keeping each result traceable and recoverable.

That meant streaming, batch, search, storage and orchestration had to work as one system.

Why it was difficult

Source formats included JATS/BITS XML and provider-specific metadata.
Processing had to remain reliable at tens of millions of events per day.
Long-running enrichment steps required replay, dead-letter handling, and idempotency.
Researchers needed full-text search and a knowledge graph from the same data.
Infrastructure needed repeatable deployment across many AWS services.
New publications had to reach researchers as they were released, not in a weekly batch.

What we delivered

Every stage retains the evidence needed to trace a derived result back to its source and to replay failed work safely.

Results

<1 min
From a new publication arriving to researchers using it
1B+
Records managed in Apache Iceberg
30M+
Events processed per day
500M+
Publication records processed, one per XML file
45M+
Entities in the knowledge graph, linked by millions of relationships
TECHNOLOGY
  • Amazon Web Services (AWS) logoAWS
  • Amazon S3 logoAmazon S3
  • Apache Iceberg logoApache Iceberg
  • Amazon MSK (Managed Streaming for Apache Kafka) logoAmazon MSK
  • AWS Glue logoAWS Glue
  • AWS Fargate logoAmazon ECS on AWS Fargate
  • AWS Lambda logoAWS Lambda
  • Amazon SQS logoAmazon SQS
  • Amazon SNS logoAmazon SNS
  • Amazon EventBridge logoAmazon EventBridge
  • Amazon OpenSearch Service logoAmazon OpenSearch Service
  • Neo4j logoNeo4j
  • Amazon DynamoDB logoAmazon DynamoDB
  • Amazon CloudWatch logoAmazon CloudWatch
  • Pulumi logoPulumi

Building a high-volume data product?

Start a Production Sprint and put your first use case into production in 6–8 weeks.