The challenge
Research publications arrived live from publishers and databases such as Springer and PubMed, in many document formats and at changing volumes. The platform had to keep the raw evidence, normalize metadata, coordinate enrichment and make every new paper usable quickly, while keeping each result traceable and recoverable.
That meant streaming, batch, search, storage and orchestration had to work as one system.
Why it was difficult
Source formats included JATS/BITS XML and provider-specific metadata.
Processing had to remain reliable at tens of millions of events per day.
Long-running enrichment steps required replay, dead-letter handling, and idempotency.
Researchers needed full-text search and a knowledge graph from the same data.
Infrastructure needed repeatable deployment across many AWS services.
New publications had to reach researchers as they were released, not in a weekly batch.
What we delivered
Every stage retains the evidence needed to trace a derived result back to its source and to replay failed work safely.
Results
<1 min
From a new publication arriving to researchers using it
1B+
Records managed in Apache Iceberg
30M+
Events processed per day
500M+
Publication records processed, one per XML file
45M+
Entities in the knowledge graph, linked by millions of relationships
TECHNOLOGY
AWS
Amazon S3
Apache Iceberg
Amazon MSK
AWS Glue
Amazon ECS on AWS Fargate
AWS Lambda
Amazon SQS
Amazon SNS
Amazon EventBridge
Amazon OpenSearch ServiceNeo4j
Amazon DynamoDB
Amazon CloudWatchPulumi