Forum Discussion
Storing and parsing HL7 data
- 17 days ago
Hi ipkus,
I would start by separating two questions: whether the source is HL7 v2 or FHIR, and whether the ingestion is batch or real-time.
For HL7 v2, my usual pattern in Fabric would be to keep the original message unchanged in a raw/Bronze layer and then create a structured analytical representation downstream.
Something like:
Raw/Bronze
-> original HL7 message + source/message metadataStructured/Silver
-> parsed clinical entities in Delta tables, or a FHIR-based canonical representationAnalytics/Gold
-> domain-specific tables for encounters, orders, labs, admissions, etc. that are convenient for semantic models and Power BII would not use a single JSON column as the only long-term analytical representation. Keeping JSON/raw text is useful for replay and auditability, but repeatedly extracting every HL7 field from that representation becomes harder to govern and query as the number of message types grows.
Microsoft's current Healthcare data solutions architecture follows a similar medallion approach. Its unified storage design includes healthcare formats such as HL7 and FHIR, while the structured healthcare model in the Silver layer is based on FHIR.
If your source is HL7 v2 and you want FHIR as the canonical model, Azure Health Data Services also provides $convert-data, which supports HL7v2 -> FHIR R4 conversion. Microsoft treats that as one component of an ETL pipeline rather than the complete ingestion solution.
One practical point for HL7 v2 is that I would use an HL7-aware parser rather than relying only on splitting by |. Real messages also contain components, repetitions, escaping and message-specific segment structures that become difficult to manage with delimiter logic alone.
For batch workloads, files can land in OneLake and be parsed/transformed with pipelines/notebooks. For real-time workloads, I would keep the same Bronze/Silver/Gold separation but change the ingestion path and make sequencing/idempotency explicit where the clinical workflow depends on message order.
There is also a current product consideration: Microsoft is changing the delivery model for Healthcare data solutions in Fabric. A customer-managed source-code package is now available, and Microsoft is transitioning away from new deployments of the current managed solution, so I would take that roadmap into account for a new implementation.
If you can share whether the source is HL7 v2 or FHIR, and whether it is batch files or a live interface feed, the recommended implementation can be narrowed down quite a bit.
AI-assisted drafting: AI was used to help structure and phrase this response. I reviewed and validated the technical content before posting.
Could you add a bit more detail about your setup? How you store HL7 for analytics can change a lot depending on whether you're working with HL7 v2 or FHIR, and whether your workflow is real‑time or batch.