Forum Discussion
Suggestions for CDC driven data ingestion
- 1 year ago
Hi PhilBrown ,
I’d encourage you to submit your detailed feedback and ideas via Microsoft's official feedback channels, such as the Microsoft Fabric Ideas.
Feedback submitted here is often reviewed by the product teams and can lead to meaningful improvement.
Thanks,
Prashanth Are
MS Fabric community support
Thanks for the reply Prashanth. A couple of notes, while I research.
In general, there always seems to be some minor to significant latency in spinning up batch jobs, often adding between 2 and 4 minutes to any notebook, pipeline execution, or Spark job. I was attempting to avoid this accumulated latency with something closer to a long-running job, but this may be better suited to Azure functions or service.
Regarding points 2 & 3, my existing spark job is already doing this effectively caching in memory and batching table upserts/deletes every 2-5 minutes.
On #4, there are config settings in Debezium that prevent full table-locks during snapshot, yet still provide acceptable concurrancy in a majority of cases. They just aren't made available within the generic EventStream setup in Fabric. Probably a big miss in my opinion.
Hi PhilBrown ,
I’d encourage you to submit your detailed feedback and ideas via Microsoft's official feedback channels, such as the Microsoft Fabric Ideas.
Feedback submitted here is often reviewed by the product teams and can lead to meaningful improvement.
Thanks,
Prashanth Are
MS Fabric community support