The challenge
Every piece of equipment on a ship arrives with its supplier’s manual: long, inconsistent documents with tables, diagrams, part references and revision history. Maintenance plans had to be found and typed into the ERP by hand, which slowed delivery and made it hard to prove which page a value came from.
Fully on-premises
Documents never leave the site
Supplier manuals contain confidential equipment data, so the model, the embeddings and the vector index all run on the customer’s own hardware. No document or prompt goes to an external AI service.
On the customer’s own GPU server
Production runs on an NVIDIA RTX PRO 6000 Blackwell GPU server in the customer’s environment. We deployed a self-hosted open-source LLM on it, so sensitive documents are processed on-site.
How we run AI on-premises →What we delivered
PDF and scanned-document processing, OCR where required, metadata capture, versioning, and durable source storage.
Content is segmented using document hierarchy and layout so tables, sections, and cross-references retain useful context.
A RAG chatbot lets engineers ask questions about any manual. Embeddings and vector retrieval find the relevant passages, and the self-hosted LLM answers with a citation to the PDF name and page.
Each supplier’s maintenance plan, with its tasks, intervals and spare parts, is extracted into a fixed schema and delivered as a structured document the ERP can import, with every value citing the PDF name and page it came from.
Every extracted field gets a confidence score. Low-confidence fields go to a person for review before anything reaches the ERP, and test sets measure extraction accuracy.
Outcome
Engineering knowledge became reusable while the original document, version, page, and extraction evidence remained available for audit and review.
Python
- OCR
- NVIDIA RTX PRO 6000 Blackwell
- Self-hosted open-source LLM


