eventhouse
18 TopicsBest pattern for keeping a current-state table in Eventhouse from Eventstream CDC?
Hi all, I'm streaming change data from Azure SQL Database into Eventhouse using an Eventstream CDC source. Ingesting the raw change events works fine. The question is how best to turn them into a queryable current-state table, the way you would with a Flink CDC or Debezium pipeline. Setup: Source: Azure SQL Database (CDC enabled), about 20 tables, mostly updates with occasional deletes Eventstream → Eventhouse, raw change events land in one table per source table Consumers: Real-Time Dashboards and some Activator rules that need the latest row per key What I'm considering: A materialized view using arg_max(timestamp, *) by key over the raw table An update policy that flattens the change envelope into a typed table, with the materialized view on top of that Handling deletes by keeping a delete flag, then filtering it out in the view or a stored function Questions: Is option 2 (update policy, then materialized view) the recommended pattern, or is there a simpler built-in approach? How are people handling deletes? Filtering the latest row by a delete flag works, but deleted keys keep taking up space. Is there a good retention or cleanup pattern? What should I use as the ordering column so out-of-order events don't produce a stale "latest" row: the source LSN/commit timestamp, or ingestion time? Any gotchas when the source schema changes (for example, a new column)? Thanks!Solved50Views1like4CommentsHandling late and out-of-order events in Eventstream windows and Eventhouse aggregations?
I'm coming from Apache Flink, where late and out-of-order events are handled with event-time watermarks and allowed lateness. I'm trying to understand the equivalent approach in Fabric Real-Time Intelligence. Scenario: IoT devices send telemetry with an event timestamp. Because of network delays, some events arrive several minutes late, and occasionally out of order. I need accurate 5-minute aggregates (average and max per device) for dashboards and Activator rules. Questions: Eventstream: when using windowed aggregations (tumbling or hopping), are windows based on event time or arrival time? Is there a setting for out-of-order tolerance or late-arrival handling, and what happens to events that arrive after a window has closed: are they dropped, or do they produce a corrected result? Eventhouse: if I aggregate in Eventhouse instead, for example with a materialized view using bin(EventTime, 5m), are late events included automatically when they arrive? Any caveats with materialized views and late data? Which layer is better for aggregations when correctness under late data matters more than lowest latency? Activator: if a rule has already fired on an aggregate that later gets corrected by late data, is there a pattern for handling that, or should rules only run on finalized windows?30Views1like2CommentsDate slicer
Hi, I need some help with Power BI. I want to use a Between Date Slicer and display the dates in this format: DD-MMM-YYYY Example: 01-Jan-2020 I don’t want to use the default dropdown/date format. I specifically want the Between slicer with two date inputs (From Date and To Date), but I need the displayed dates to appear as 01-Jan-2020 instead of the default format. Could you please guide me on how to achieve this? If the standard Power BI slicer does not support this format, is there any custom visual or alternative solution that can provide the same Between Date functionality with the DD-MMM-YYYY format? Thanks!Solved59Views1like5CommentsFabric Business Events: what delivery guarantees and replay pattern should we design for?
Hi all, I am testing the newer Business Events capability in Fabric Real-Time Intelligence and trying to understand what reliability assumptions should be made for a production design. The pattern I am looking at is roughly: Eventstream → Business Event → Activator → downstream action / User Data Function with Eventhouse enabled so the published business events are also retained for historical analysis. The current documentation explains the publisher/consumer model and shows how Eventstream can publish a governed business event that Activator then consumes. What I have not been able to find clearly documented is the delivery contract between the published business event and its consumers. A few things I am trying to clarify: If an Activator consumer or downstream action is temporarily unavailable, does Fabric retry delivery of the business event? Should consumers assume at-least-once delivery and therefore be designed to handle duplicate events, or is a different delivery model used? Is event ordering guaranteed in any scope, for example for events from the same Eventstream publisher? Since published business events can also be retained automatically in Eventhouse, is that retained history intended to support replay/reprocessing after a consumer outage, or is it primarily an analytical record and replay would need to be implemented separately? Are there documented retry or delivery-retention windows that should be considered when designing an operational workflow? I am mainly trying to understand what a resilient production pattern should look like when the business event triggers something with side effects, where processing the same event twice or silently missing an event would matter. Would you generally make the downstream consumer idempotent and treat Eventhouse as an audit/recovery store, or is there a more Fabric-native pattern for this? Interested to hear how others are approaching this with Business Events and Activator.121Views0likes3CommentsBest practice for handling schema evolution in Fabric Eventstream before data reaches Eventhouse?
I have an Eventstream receiving operational events where the schema may evolve over time. For example, the producer initially sends: DeviceId, Timestamp, Temperature, Status but later adds fields such as: Location, FirmwareVersion, ErrorCode I want the pipeline to continue ingesting events without breaking downstream KQL tables, update policies, materialized views, or Real-Time Dashboards. I am trying to understand where schema evolution should ideally be handled in a production Fabric RTI architecture. Would you: enforce the contract upstream using Schema Registry normalize changing fields inside Eventstream before Eventhouse ingestion land the raw payload first and handle schema evolution inside Eventhouse/KQL maintain separate versioned event schemas/tables How are people handling this in production when producers can add fields without notice? I am particularly interested in avoiding a design where every small upstream schema change forces updates across Eventstream, KQL tables, update policies, and downstream dashboards.Solved126Views0likes2CommentsBest practice for deciding between Eventstream transformations and Eventhouse update policies
Hi Fabric Community, I am exploring a Real-Time Intelligence architecture and would appreciate some guidance on where transformation logic should ideally be placed. The proposed flow is: Azure Event Hubs → Fabric Eventstream → Eventhouse → Real-Time Dashboard / Power BI Fabric Eventstream supports filtering, field management, aggregation and other processing before events are written to the destination. An Eventhouse can also ingest the raw events first and transform them into curated tables through KQL update policies. I am trying to understand the recommended boundary between these two layers. For example, assume the incoming event contains: Device or customer identifier Event timestamp Event type Location Numeric readings Additional JSON properties The required processing includes: Removing events that fail basic validation Renaming and standardizing fields Converting timestamps and data types Flattening selected JSON properties Enriching the event with reference data Creating five-minute aggregates Preserving the original event for auditing and future reprocessing My current thinking is: Use Eventstream for lightweight filtering, routing and simple schema normalization. Land the original event in a Bronze table whenever replay or auditing is required. Use Eventhouse update policies or KQL for enrichment, reusable business logic and curated Silver tables. Use materialized views for frequently queried aggregations rather than calculating them repeatedly in dashboards. However, I am unsure where Microsoft recommends drawing the line. A few questions: Are there transformation types that should generally remain in Eventstream rather than Eventhouse? Is it considered good practice to send both the raw stream and a transformed derived stream into separate Eventhouse tables? When using Eventstream’s Event processing before ingestion mode, what are the trade-offs compared with direct ingestion followed by an Eventhouse update policy? How do teams handle changes to transformation logic when historical events need to be reprocessed? For reference-data enrichment, would you normally perform the lookup in Eventstream or after ingestion with KQL? Are five-minute or hourly aggregations better implemented in Eventstream, through an update policy, or with an Eventhouse materialized view? Microsoft’s Eventstream destination guidance documents both direct ingestion and event processing before ingestion, while the KQL update-policy documentation provides another way to transform ingested data. I would be interested to hear how others divide responsibility between Eventstream and Eventhouse in production, particularly where auditability, reprocessing and maintainability are important. Thanks in advance!Solved124Views0likes2CommentsHow do you decide when to use Real-Time Intelligence instead of batch processing in Microsoft Fabric
Hi everyone, I'm learning Microsoft Fabric and recently started exploring Real-Time Intelligence. I understand that it can process streaming data, but I'm trying to understand when it's the right choice compared to traditional batch processing. I have a few questions: What types of business scenarios benefit the most from Real-Time Intelligence? When would you choose streaming over scheduled batch processing? What are some common real-world use cases you've worked on? Are there any performance or cost considerations that beginners should be aware of? I'd really appreciate hearing about your experiences and any best practices you recommend. Thank you!Solved379Views1like6CommentsReal‑Time Dashboards cannot be created
Hi all. We’re currently facing an issue in the a Fabric tenant that prevents us from creating Real‑Time Dashboards. According to Microsoft documentation, this feature requires either a Fabric Premium capacity or a Premium Per User (PPU) license. Reference: Create a Real-Time Dashboard - Microsoft Fabric | Microsoft Learn For example: In one tenant (F4 capacity, North Europe) with a user licensed as Premium Per User, the feature works as expected. In another tenant (F16 capacity, Southeast Asia) with a user licensed as Pro, Real‑Time Dashboards cannot be created. The error shown is: “Real‑Time Dashboard item can only be created if this action is enabled in the tenant. Ask your tenant admin to enable this.” This error is confusing, as the option to enable the feature is no longer available in the Admin Portal. Anyone had a similar issue?6.6KViews2likes15CommentsWorking with Microsoft Fabric Support - A Collaborative Approach
At Microsoft Fabric Support, our goal is simple: help you resolve issues as quickly and smoothly as possible. Just like the recommendations shared in the Microsoft Fabric Community’s guidance on getting questions answered effectively, providing clear and complete information upfront helps everyone move faster and avoid unnecessary back-and-forth. Support is most effective when it’s a collaborative process. You bring the knowledge of your environment and workloads, and we bring deep platform expertise and diagnostics to help investigate the issue together. Help Us Help You When opening a support case, please include as much of the following information as possible: Issue Details Clear description of the problem Expected behavior vs. actual behavior Is the issue intermittent or consistently reproducible? Approximate timeframe of when the issue occurred (UTC preferred) Error Information Please include: Full error message text Screenshots (highly recommended) Activity IDs / Correlation IDs if available Even small details can significantly speed up investigation. Environment Information That Helps Investigation Depending on the Fabric workload being used, please provide the relevant item IDs and environment details. Workspace Information Workspace ID / Workspace URL To find this: Open the Fabric workspace Copy the URL from your browser Example: ".. https://app.fabric.microsoft.com/groups/<WorkspaceID>/… …" The value after /groups/ is the Workspace ID. Screenshot example above showcases the WorkspaceID and EventhouseID. Fabric Item IDs Please provide the item ID related to the affected workload, such as: EventhouseID ActivatorID LakehouseID WarehouseID SemanticModelID EventstreamID PipelineID These IDs help us locate telemetry and backend diagnostics more efficiently. You can typically find these: In the browser URL while inside the Fabric item Within item settings/details pages From the Fabric portal navigation pane Additional Environment Details Please also include: "… WorkspaceID: Capacity Name: Region …" Example: "… WorkspaceID: xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx Capacity Name: Fabric-Prod-EastUS Region: East US …" This information helps us identify: Capacity-level issues Regional service impact Configuration-specific behavior Why This Matters Microsoft Fabric is a distributed cloud platform, and troubleshooting often depends on correlating logs, telemetry, timestamps, and resource identifiers together. Providing complete details upfront helps: Reduce delays Avoid repeated clarification requests Accelerate root cause analysis Create a smoother support experience for everyone involved We’re here to partner with you throughout the process, and we truly appreciate your collaboration in helping us investigate issues effectively. We look forward to working with you!1.1KViews5likes3CommentsDigital Twin Builder (preview): Bounding storage growth for an incremental time-series
Is there any supported way to apply a retention horizon to a Digital Twin Builder time-series mapping that uses incremental processing, so the underlying dtdm table doesn't grow without bound? My scenario: high-frequency live telemetry landing in an Eventhouse, surfaced to DTB through a lakehouse table (I've used both shortcuts and pipeline copy activities in my tests). The static layer and relationships are working well. The question is purely about the time-series layer. With incremental mapping enabled, the series appends continuously and the dtdm table grows linearly with no eviction. I've confirmed a couple of things that don't solve it, so I'll pre-empt them: Retention on the source table doesn't bound the twin copy. Incremental processing is forward-only, so source rows aging out are never removed from data the mapping already ingested. A mapping filter governs ingestion, not eviction, and only limits initial backfill (no relative/dynamic date filter). I assume full reload (incremental disabled) over a windowed source table does keep things bounded, since each run replaces the prior load, and that's my current fallback. But I'd like to confirm whether there's any retention/TTL control I've missed for the incremental path specifically, at the mapping, entity, or item level, or any sanctioned way to age out rows from the dtdm time-series table. If the answer is "not currently supported," that's useful to know too, and I'll file it as an Idea. Running on a Fabric capacity, DTB still in preview. Thanks in advanceSolved1.1KViews0likes5Comments