eventhouse
80 TopicsBest pattern for keeping a current-state table in Eventhouse from Eventstream CDC?
Hi all, I'm streaming change data from Azure SQL Database into Eventhouse using an Eventstream CDC source. Ingesting the raw change events works fine. The question is how best to turn them into a queryable current-state table, the way you would with a Flink CDC or Debezium pipeline. Setup: Source: Azure SQL Database (CDC enabled), about 20 tables, mostly updates with occasional deletes Eventstream → Eventhouse, raw change events land in one table per source table Consumers: Real-Time Dashboards and some Activator rules that need the latest row per key What I'm considering: A materialized view using arg_max(timestamp, *) by key over the raw table An update policy that flattens the change envelope into a typed table, with the materialized view on top of that Handling deletes by keeping a delete flag, then filtering it out in the view or a stored function Questions: Is option 2 (update policy, then materialized view) the recommended pattern, or is there a simpler built-in approach? How are people handling deletes? Filtering the latest row by a delete flag works, but deleted keys keep taking up space. Is there a good retention or cleanup pattern? What should I use as the ordering column so out-of-order events don't produce a stale "latest" row: the source LSN/commit timestamp, or ingestion time? Any gotchas when the source schema changes (for example, a new column)? Thanks!Solved50Views1like4CommentsHandling late and out-of-order events in Eventstream windows and Eventhouse aggregations?
I'm coming from Apache Flink, where late and out-of-order events are handled with event-time watermarks and allowed lateness. I'm trying to understand the equivalent approach in Fabric Real-Time Intelligence. Scenario: IoT devices send telemetry with an event timestamp. Because of network delays, some events arrive several minutes late, and occasionally out of order. I need accurate 5-minute aggregates (average and max per device) for dashboards and Activator rules. Questions: Eventstream: when using windowed aggregations (tumbling or hopping), are windows based on event time or arrival time? Is there a setting for out-of-order tolerance or late-arrival handling, and what happens to events that arrive after a window has closed: are they dropped, or do they produce a corrected result? Eventhouse: if I aggregate in Eventhouse instead, for example with a materialized view using bin(EventTime, 5m), are late events included automatically when they arrive? Any caveats with materialized views and late data? Which layer is better for aggregations when correctness under late data matters more than lowest latency? Activator: if a rule has already fired on an aggregate that later gets corrected by late data, is there a pattern for handling that, or should rules only run on finalized windows?30Views1like2CommentsChanging Ingestions from JSON
I have an EventStream receiving IoT Data out of JSONs. In my KQL table the columns timestamp, deviceId and for example x and y exist. The data from one sensor gets ingested in the columns timestamp, deviceId and y. No x since there is no data in the JSON for that. I dont work with a Mapping because it only makes it more complicated and the names in the JSON and in the KQL table are identical. No I changed the named for data from y to x when creating the JSON. Since there is no mapping and the column already exists in the KQL table the data should be saved now in column x. My problem is, it doesn't the data is not in x nor in y. It seems to be gone. Is that a bug in fabric? Can I solve this?129Views0likes11CommentsDate slicer
Hi, I need some help with Power BI. I want to use a Between Date Slicer and display the dates in this format: DD-MMM-YYYY Example: 01-Jan-2020 I don’t want to use the default dropdown/date format. I specifically want the Between slicer with two date inputs (From Date and To Date), but I need the displayed dates to appear as 01-Jan-2020 instead of the default format. Could you please guide me on how to achieve this? If the standard Power BI slicer does not support this format, is there any custom visual or alternative solution that can provide the same Between Date functionality with the DD-MMM-YYYY format? Thanks!Solved59Views1like5CommentsGlobal Aircraft✈️ Live Tracking with Microsoft Fabric Real-Time Intelligence
In this blog, I’ll walk you through, build a real-time global flight tracking system. We will ingest live flight data from the public OpenSky Network API using a Python polling script, stream it through Microsoft Fabric Eventstream into an Eventhouse (KQL Database), transform the dense raw arrays using KQL update policies and visualize the results on a Real-Time Dashboard complete with maps, KPIs and analytical charts. Prerequisites: Valid Fabric Capacity / Trail License Knowledge on Python Knowledge on KQL Step1: Setup Workspace & Eventhouse (KQL Database) Created a workspace “FlightTracking-[WS]” Created an eventhouse “FLightTracking-EH” Create a raw ingestion table “RawFlightbatc” in KQL Databse Step 2: Setup a Fabric Eventstream Created a Eventstream “GlobalFlightStream” and select ‘Use custom endpoint’ Click ‘Add’ Click on ‘Publish’ Copy the ‘Event hub name’ and ‘connection string-primary key’ into notepad Step 3: Notebook Creation and Setup Python script Created a notebook. Make sure select Python Install azure eventhub package Let’s go back to eventstream and add destination by selecting ‘Eventhouse’ Configure all details and click on save Comeback to Notebook and insert the Python script which is having all the connection strings / passwords etc. Run the notebook Now notebook started running and sending the data to eventstream Data is loaded into eventstream and Click on Publish Now eventstream is ‘Live’ Data is loading into KQL Database Step 4: Regularizing Data with KQL & Update Policies Create a cleaned table “FlightStates” Create the parsing function and update policy Alter table with updated policy Step 5: Building Real-Time Dashboard Click on Realtime dashboard and give a name “FlightOperationsDashboard” Now click on edit Run the below code to get total active flights count Change the chart to Stat and rename, click on Apply KPI added and click on Add visual and take new Stat visual Insert the code and run to get India Origin Flights and Format it and click Apply In the same way, I built other KPIs Select Map Chart Run the below code and Fill all details and Click on Apply. Here we’re calculating Flights trend We can see chart added to Dashboard Select a Bar chart Run the below code, fill all details and click on apply to add Bar chart to Dashboard. Here we’re getting the top 10 countries by aircrafts Run the below code, fill all details and click on apply to add Column chart to Dashboard. Here we’re categorizing the baro-altitude which is critical for airport delay predection Run the below code, fill all details and click on apply to add Pie chart to Dashboard. It splits the aircraft parked versus those actively flying Together, these visuals transform raw flight telemetry into an operational monitoring experience. Key takeaways: Fabric Eventstream provides a streamlined way to ingest external real-time feeds. Eventhouse/KQL Database provides a real-time analytical environment for flight telemetry. KQL mv-expand simplifies the processing of nested flight-state arrays. Update Policies automate transformation from raw streaming data into structured analytical data. KQL geospatial functions enable location-based flight analysis. Real-Time Dashboards transform streaming telemetry into actionable operational insights. Conclusion: This project demonstrates how Microsoft Fabric Real-Time Intelligence can be used to build an end-to-end real-time aircraft tracking solution. we can transform continuously arriving aircraft telemetry into meaningful real-time insights. Do you want to replicate? Get Code file from my GitHub link You can find all the KQL queries, update policy functions, and the complete Python polling script in the official GitHub repository below: [Download from here] Acknowledgements I would like to express my sincere gratitude to @SuryaTejaJosyul , @minniwalia and @rajendraongole1 for their continuous guidance and support throughout this Real-Time Intelligence (RTI) implementation. Their insights and encouragement played a key role in helping me complete this solution successfull Happy learning! — Inturi Suparna Babu [LinkedIn]Fabric Eventstream MySQL CDC scans all databases and does not emit row-level events
Hello Fabric Team, We are testing the Microsoft Fabric Eventstream MySQL CDC connector with an AWS RDS MySQL database. The MySQL instance is accessible through a private IP, and Fabric connects through a VNet data gateway/private network connection. Connectivity is successful, and the connector remains Active without reporting processing errors. Our MySQL environment contains the same table structure across multiple databases, following a pattern similar to: application_<tenant>.sample_table We configured the source to capture only one fully qualified table: application_tenant1.sample_table The MySQL configuration is: log_bin = ON binlog_format = ROW binlog_row_image = FULL The connection user has SELECT, SHOW DATABASES, REPLICATION CLIENT, REPLICATION SLAVE/REPLICA, and the required snapshot permissions. Each Fabric source uses a unique Server ID. CDC from the same RDS instance has previously worked successfully through another CDC platform. However, we are seeing the following behavior: Although only one fully qualified table is selected, Fabric/Debezium scans schema metadata for every accessible database and table on the MySQL server. More than 45,000 schema events were generated across the tenant and history databases. The schema events contain DDL such as DROP TABLE IF EXISTS and CREATE TABLE. Snapshot mode is configured as Initial. The selected table contains approximately 16,000 records, but Fabric emitted zero snapshot r events. We committed a new insert into the selected table after schema discovery finished, but no c event was emitted. Eventhouse receives the schema events successfully, confirming that the source-to-Eventhouse path is operational. The Eventhouse destination stores the Debezium payload in a dynamic column using a mapper from payload to RawEvent. Could you please confirm the following? Is it expected for the Fabric-managed MySQL CDC connector to scan and publish schema events for every database when only one fully qualified table is selected? Does Fabric expose or support settings equivalent to: database.include.list table.include.list schema.history.internal.store.only.captured.databases.ddl schema.history.internal.store.only.captured.tables.ddl Why would the connector finish schema discovery but not emit existing rows from the selected table when Snapshot mode is Initial? Why are committed inserts not emitted after the schema phase appears to finish? Is there a known limitation involving AWS RDS MySQL accessed through a private IP and VNet data gateway? What is the recommended supported destination pattern for this CDC stream? Our main objective is reliable ongoing CDC capture. We are flexible about the landing destination and can use: Eventhouse ADLS Gen2 in JSON or Parquet Fabric Lakehouse/OneLake Snowflake as the final destination Please let us know which logs, connector identifiers, timestamps, or additional configuration details would help investigate this behavior. Thank you.Solved63Views0likes4CommentsUnable to create a Monitoring Eventhouse in a Microsoft Fabric workspace.
I am trying to enable Workspace Monitoring for a Microsoft Fabric workspace. The user has Fabric Administrator access and is also a Workspace Administrator. However, under: Workspace Settings → Monitoring the + Eventhouse option is disabled/not available. We have verified the tenant-level settings and confirmed that “Workspace admins can turn on monitoring for their workspaces” is enabled.Solved131Views2likes6CommentsAnomaly Detector is disabled
Hello, I have a big Table in which we have our TimeSeries data from various devices. Since not every device sends the same data we have a big schema of numeric columns that are not filled for every deviceId. Is that the reason why there is no Anomaly Detector in my EventHouse? In the schema I can see that columns for the data are int or real. Some columns are string, but that shoudn't be the reason, right? Has anyone faced the same issue?93Views0likes5CommentsFabric Business Events: what delivery guarantees and replay pattern should we design for?
Hi all, I am testing the newer Business Events capability in Fabric Real-Time Intelligence and trying to understand what reliability assumptions should be made for a production design. The pattern I am looking at is roughly: Eventstream → Business Event → Activator → downstream action / User Data Function with Eventhouse enabled so the published business events are also retained for historical analysis. The current documentation explains the publisher/consumer model and shows how Eventstream can publish a governed business event that Activator then consumes. What I have not been able to find clearly documented is the delivery contract between the published business event and its consumers. A few things I am trying to clarify: If an Activator consumer or downstream action is temporarily unavailable, does Fabric retry delivery of the business event? Should consumers assume at-least-once delivery and therefore be designed to handle duplicate events, or is a different delivery model used? Is event ordering guaranteed in any scope, for example for events from the same Eventstream publisher? Since published business events can also be retained automatically in Eventhouse, is that retained history intended to support replay/reprocessing after a consumer outage, or is it primarily an analytical record and replay would need to be implemented separately? Are there documented retry or delivery-retention windows that should be considered when designing an operational workflow? I am mainly trying to understand what a resilient production pattern should look like when the business event triggers something with side effects, where processing the same event twice or silently missing an event would matter. Would you generally make the downstream consumer idempotent and treat Eventhouse as an audit/recovery store, or is there a more Fabric-native pattern for this? Interested to hear how others are approaching this with Business Events and Activator.121Views0likes3CommentsEventhouse Capacity Planner minimum CU not reflected in UI/API
Hi all, We've set a minimum of 32 CU on our Eventhouse via Capacity Planner (autoscale alone isn't sufficient - we need guaranteed baseline capacity to protect a large bulk-ingestion workload from destination-side OutOfMemory during a migration). We've been told the 32 CU setting has been applied on the backend, but the Fabric UI and REST API (.show cluster / .show diagnostics) still report capacity/behavior consistent with a much lower tier. Can anyone confirm: Whether the a custom minimum CU value is actually enforced server-side even though the UI/API don't show it, and When UI/API reporting is expected to catch up to reflect the configured minimum? Any insight or similar experience would be appreciated.Solved148Views0likes6Comments