Who uses Soria
Data Engineers
Build and manage ingestion pipelines via the MCP server in Claude Code. Define scrapers, configure schemas, run extraction, and publish warehouse models.
Analysts
Explore structured healthcare datasets through dashboards, filter data, search across all datasets and news stories, and query the warehouse directly.
Core capabilities
AI Extraction
Gemini-powered detection and extraction handles PDFs, Excel, and CSVs — no manual template authoring required.
Data Warehouse
A four-layer pipeline (bronze → silver → gold → platinum) transforms raw files into dashboard-ready views.
MCP Integration
Connect Claude Code or any MCP-compatible AI assistant to run the full pipeline with natural language.
News Intelligence
An automated daily pipeline fetches, scores, clusters, and summarizes healthcare news from configured sources.
Two pipeline paths
Soria handles two types of source files differently: PDFs and Excel (unstructured → structured): Soria detects which pages contain relevant data, extracts it into CSVs using Gemini against your defined schema, validates accuracy, and normalizes inconsistent values. CSVs (already structured): Soria maps source column headers to your canonical schema columns, handling schema drift when column names change across files or over time. Both paths land data in the warehouse for SQL modeling and dashboarding.Next steps
Quick Start
Connect via MCP and run your first command in minutes
Pipeline Overview
Understand the end-to-end data pipeline