data
9 TopicsMaking AI Work: Aligning Data, Teams, and Business Goals
AI adoption often stalls not due to a lack of tools, but because of what I call “frozen yoghurt syndrome”—an overwhelming number of choices that leads to confusion, misalignment, and stalled progress. This session offers a pragmatic framework for building a scalable data strategy that cuts through the noise and focuses on delivering business value. We’ll explore four key pillars: - Identify the Problem and Bigger Picture – Start with the business need. AI isn’t always the right answer; sometimes automation or simpler analytics are more effective. - Design the Strategy Around Data, AI, and the Use Case Together – Align initiatives with clear business goals and define success metrics that matter to both technical and non-technical teams. - Build the Foundation – Encourage cross-functional collaboration, select scalable tools, and document everything from data sources to model assumptions. - Create a Positive Feedback Loop Between Tech and Non-Tech Teams – Foster a shared product language, enable self-service access to data, and build trust through transparency and iteration. Attendees will leave with practical frameworks to ensure AI efforts are not just technically sound — but strategically aligned, collaborative, and sustainable.126Views0likes0CommentsAll About Data Retention: Tactics for optimizing computing resources
Topic: All About Data Retention: Tactics for optimizing computing resources Abstract: How long should data be kept for? How should it be handled as it moves from actively used hot data to seldom used cool data? What tactics can be used to optimize computing resources when handling data of different types, from different systems, and of different ages. Retention is a seldom-discussed topic, but one that can be used to reduce cloud hosting bills, improve performance, improve predictability, and ensure that an application remains in compliance with legal and contractual obligations, This session dives into data retention from all angles and provides a wide variety of tactics for creating retention policies and implementing technical solutions. Speaker: Edward Pollack, Data Architect & Microsoft Data Platform MVP Speaker Profile: Ed Pollack is a Microsoft Data Platform MVP with a passion for learning how Data Platforms work and sharing that knowledge with the community. His experiences in data architecture, database design, performance optimization, and data security are motivation for public speaking, writing, coding, and other community activities. Ed has spoken at SQL Saturday events, SQL Bits, PASS Summit, EightKB, and many other regional and international events. Ed is the organizer of the Capital Area SQL Server Group and SQL Saturday Albany, as well as a co-organizer of SQL Saturday New York City, and Future Data Driven. He has published a number of books, including "Dynamic SQL: Applications, Performance, and Security in Microsoft SQL Server", "Expert Performance Indexing in Azure SQL and SQL Server 2022", and "Analytics Optimization with Columnstore Indexes in Microsoft SQL Server: Optimizing OLAP Workloads". Ed is also an active contributor of content to SimpleTalk. In his free time, Ed enjoys video games, traveling, cooking exceptionally spicy foods, and hanging out with his amazing wife and sons.47Views0likes0CommentsMaximize SQL Server Performance with Read Committed Snapshot Isolation
Data Engineers in Toronto May 2026 Semimonthly Meeting Topic: Maximize SQL Server Performance with Read Committed Snapshot Isolation Abstract: Are your read operations frequently blocked by write operations? Do you want to retrieve data without relying on NOLOCK? Then it's time to switch to Read Committed Snapshot Isolation (RCSI). In this session, I’ll explain what RCSI is, how it works, and how it differs from the default Read Committed Isolation level. I’ll demonstrate how version store comes into play with this isolation level and when it can overwhelm TempDB if not managed properly. Finally, I’ll discuss how to manage the implications of enabling RCSI. By the end of this session, you'll have a clear understanding of RCSI and when to implement it in your environment for improved concurrency and performance. Speaker: Haripriya Naidu, Lead SQL Server Database Administrator Speaker Profile: Haripriya Naidu has been working as a SQL Server Database Administrator for 11 years in Boston, MA. She is passionate about databases and enjoys diving into SQL Server engine internals and optimizing performance. She is an AWS Certified Solutions Associate and won "Rookie of the Year" Award at work in 2024 for outstanding performance and "Bookworm of the Year" Award at work in 2025 for continuously learning and sharing knowledge with the team. She recently started her speaking journey with New Stars of Data and spoke at SQL Saturdays Boston, Albany and San Diego and PASS Summit 2024. The meeting is over Microsoft Teams, and the joining link is https://teams.microsoft.com/l/meetup-join/19%3ameeting_NzZkYWIyOTAtODk1MC00MjVmLWJlNjUtNTRiODZmODA2Zjdh%40thread.v2/0?context=%7b%22Tid%22%3a%22bd9727e8-f539-4c76-983c-6c30130c0bee%22%2c%22Oid%22%3a%229e8d5a64-e773-4ca2-90f6-9a266129171e%22%7d See you at the meeting!61Views0likes0CommentsData Modernization at Traditional Banks - From Legacy Mainframe to Cloud
(Date updated to 1 June 2026) Data Engineers in Toronto May 2026 Semimonthly Meeting Topic: Data Modernization at Traditional Banks - From Legacy Mainframe to Cloud Abstract: As technology evolved from the age of industrialization to the age of information, and finally to this current age of AI, the average lifecycle of technology shrank from 10 to 6, and now it is 3 years. Therefore, the definition of “legacy” has shifted. Applications developed before 2010, especially those built on mainframes or outdated technologies, are increasingly seen as candidates for modernization. Mainframe transformation has been a top priority in financial services because of its operational inefficiencies and high maintenance costs associated with mainframes when compared to modern distributed systems. Still, the majority of the core processing engines within financial services depend on Mainframe as the mission-critical business rules are embedded Speaker: Vishal Sharma, VP - Software Engineering at Broadridge Speaker Profile: As a Software Engineering leader, I bring over 20 years of expertise in enterprise architecture and digital transformation. By leading cross-functional teams, I enable the delivery of secure, scalable solutions that optimize operations and enhance client engagement. My efforts focus on integrating innovative technologies to align with strategic business goals, resulting in measurable outcomes for stakeholders. With a strong foundation in business architecture and fintech innovation, I excel in modernizing platforms to meet evolving industry demands. Equipped with certifications such as TOGAF, I prioritize seamless system integration, cost efficiency, and regulatory alignment to drive long-term value for organizations. Passionate about solving complex challenges, I work collaboratively to empower teams and achieve sustainable success. The meeting is over Microsoft Teams, and the joining link is https://teams.microsoft.com/l/meetup-join/19:meeting_NzZkYWIyOTAtODk1MC00MjVmLWJlNjUtNTRiODZmODA2Zjdh@thread.v2/0?context={"Tid":"bd9727e8-f539-4c76-983c-6c30130c0bee","Oid":"9e8d5a64-e773-4ca2-90f6-9a266129171e"} See you at the meeting!82Views0likes0CommentsData Mesh as the Foundation for AI/ML in Financial Services
Data Engineers in Toronto June 2026 Semimonthly Meeting Topic: Data Mesh as the Foundation for AI/ML in Financial Services Abstract: Financial institutions want AI/ML at scale, but brittle data pipelines, silos, and compliance demands slow progress. This talk shows how a Data Mesh—domain-oriented ownership, data-as-product, self-serve platforms, and federated governance—becomes the foundation for reliable, reusable ML features and trustworthy models. We’ll map mesh principles to FS use cases—fraud detection, risk, personalization—and show patterns for feature stores, lineage, quality, and access controls that satisfy regulators while accelerating delivery. Attendees will get a pragmatic blueprint: where to start, how to sequence capabilities, metrics that prove value, and pitfalls to avoid on the road from pilots to production. Speaker: Santosh Durgam, Data Engineering & Analytics Leader Speaker Profile: Santosh Durgam is a data engineering & analytics leader with 20+ years building governed, high-scale data platforms across retirement/401(k), broader financial services, and healthcare. He leads cross-functional teams that deliver production-grade data lakes, lineage-aware pipelines, and ML-enabled analytics on cloud—translating governance into measurable business outcomes. Recent speaking includes SQL Saturday Minnesota 2025, where he presented “From Ingestion to Insights: Building Robust Data Pipelines in AWS” to an in-person community audience. He has also contributed to international research forums and science conferences, and is invited to speak at ICDPN-2025 (International Conference on Data Processing & Networking), engaging practitioners and scholars on data engineering, governance, and analytics at scale. Santosh actively publishes/curates work via Google Scholar and shares practical playbooks for data quality, metadata/lineage, and operating models that connect data platforms to financial decisioning. Beyond delivery, Santosh serves the community as a peer reviewer of scholarly work on data/ML methodologies and as a judge/mentor for select industry and academic competitions, reinforcing peer validation and public recognition. He champions modern data culture—mentoring engineers and product leaders, and advocating automation (incl. AI agents) to elevate reliability, speed, and auditability in regulated environments. Santosh recently completed his Executive MBA, sharpening strategy and value-creation at the intersection of data, risk, and growth The meeting is over Microsoft Teams, and the joining link is https://teams.microsoft.com/l/meetup-join/19%3ameeting_NzZkYWIyOTAtODk1MC00MjVmLWJlNjUtNTRiODZmODA2Zjdh%40thread.v2/0?context=%7b%22Tid%22%3a%22bd9727e8-f539-4c76-983c-6c30130c0bee%22%2c%22Oid%22%3a%229e8d5a64-e773-4ca2-90f6-9a266129171e%22%7d See you at the meeting!87Views0likes0CommentsIntroduction to PySpark in Microsoft Fabric
Data Engineers in Toronto July 2026 Semimonthly Meeting Topic: Introduction to PySpark in Microsoft Fabric Abstract: With all of the engineering features in Microsoft Fabric, which medium should you use to move and transform data? Low-code data flows and pipelines? Good old relational SQL? What about this newfangled PySpark everyone is buzzing about? If the last option piques your curiosity and you haven't tried it, this is the session for you. I'll cover basic Python principles that will make even complicated Python easy to read. Then I will explain important topics to understand when managing your Spark environment. Finally, I'll showcase Fabric features and community content that can support your next steps in learning to implement PySpark in Fabric. Speaker: Jared Kuehn, Data Engineer Speaker Profile: For over a decade, Jared has been a data engineering consultant implementing Microsoft products. For more than three decades, he has been honing his skills in theater and other performing arts. As a speaker, Jared marries these two disciplines together, creating dynamic presentations that add entertainment to education. This symbiotic relationship can improve presentation engagement, support attendees in knowledge retention, and foster a culture of passion for the data industry. As a speaker, Jared has spoken at events such as: -Fabcon Vegas (Microsoft Fabric Community Conference) -DataCon Seattle (Microsoft Data Conference) -PASS Summit -SQL Saturdays, Multiple -Microsoft Fabric Global Online Conference -Future Data Driven Summit -GroupBy Conference -and more! He also runs the YouTube channel DataBard, focused on making Data fun and teaching techniques along the way. Technically-speaking, Jared has multiple Microsoft certifications in both on-prem and cloud technologies. He has extensive knowledge in: -Microsoft Fabric -SQL Query Development and Performance tuning -Kimball Data Modeling -ETL design patterns -Azure SQL DB and Azure SQL MI In his spare time, he continues to perform in community theater productions, as well as musically for his church. Upon request, he would be happy to assist with more theatrical portions of events, such as performing musical numbers. See you at the meeting!65Views0likes0CommentsStop Guessing: Solve SQL Performance Problems with Query Store
Data Engineers in Toronto July 2026 Semimonthly Meeting Topic: Stop Guessing: Solve SQL Performance Problems with Query Store Abstract: When query performance suddenly changes, finding the root cause isn't always easy. Execution plans can change over time, and the plan cache often provides only a limited view of what happened. Query Store helps by capturing query history, execution plans, and runtime statistics, giving you the visibility needed to identify regressions, compare plans, and maintain consistent performance. In this session, you'll learn how to use Query Store to monitor query behavior, investigate performance issues, manage execution plans, and leverage built-in tuning capabilities in SQL Server and Azure SQL Database. Key topics: Query Store fundamentals Query and plan history analysis Performance regression detection Query Store reports and insights Plan forcing and plan management Automatic tuning and best practices Through live demos, you'll see how Query Store makes performance troubleshooting faster, simpler, and more reliable. Speaker: Deepthi Goguri, SQL Database Administrator Speaker Profile: Deepthi is a SQL Server Database Administrator with several years of experience in Administering SQL Servers. She is a Microsoft Data Platform MVP, Microsoft certified trainer and Microsoft certified professional with an Associate and Expert level Certification in Data Management and Analytics. Deepthi blogs for DBANuggets.com. Deepthi is a Co-Organizer for Microsoft Data and AI South Florida user group, Data TGIF, Cloud Data Driven User Group, Future Data Driven Summit, Databash Conference and Data Platform Diversity, Equity, and Inclusion Virtual Group. She is also DEI Steering Committee member for PASS Data Community Summit. She is a Redgate Community Ambassador. Along with this, Deepthi loves arts and crafts. You can contact her on Twitter @dbanuggets.118Views0likes0CommentsRunning GitHub Actions Offline: Debug CI/CD workflows Locally
Data Engineers in Toronto August 2026 Semimonthly Meeting Topic: Running GitHub Actions Offline: Debug CI/CD workflows right from your local machine Abstract: Ever wanted to try out your GitHub Actions without waiting for the cloud to spin up the runners? Well, you can! Testing GitHub Actions locally lets you build, test, and debug your CI/CD workflows right from your machine. Through a tool such as act, you can simulate GitHub’s environment for your workflows locally using Docker. This means no more pushes of dozens of commits just to see if your syntax in the YAML file is wrong. You can make changes, execute, and validate your automation as many times as you want, even on a plane where there's no Wi-Fi. Just install Docker, install act, load your secrets locally, and run commands like act push to test your events in the workflows. That's faster, lightweight, and the greatest productivity booster. Taking GitHub Actions offline is not about isolation; it's about freedom and control over your automation pipeline. Speaker: Steve Yonkeu, Software Engineer & Microsoft MVP Speaker Profile: Steve Yonkeu is a backend software engineer dedicated to delivering excellence in architecture, performance, and modularity. For over 5 years he has been leading teams and projects and open source communities worldwide. Steve is also the founder of Django Cameroon and Python Cameroon communities today. He is also an organizer at PyCon Africa and speaker at PyConUS.98Views0likes0CommentsStreaming your database - Easier said than done?
Data Engineers in Toronto September 2026 Semimonthly Meeting Topic: Streaming your database - Easier said than done? Abstract: Do you need to stream your database? So did I! Here’s what to expect. In my journey as a contractor for a major cybersecurity company, I implemented a Change Data Capture (CDC) system that aimed to seamlessly stream client databases into a centralized cloud solution for analytics and enhanced features. However, after deploying the code, I quickly realized that the reality was far more complex than anticipated. Despite using an off-the-shelf product, I encountered unexpected challenges related to scaling, data formats, integrations, and deployment—issues that took me by surprise and required significant adjustments. Over the course of 18 months, my team and I navigated these hurdles to create a robust system designed for scalability and resilience against failures. In this talk, I’ll share key lessons learned and practical tips for developers and architects embarking on similar projects. You'll gain insights into preparing for potential surprises and overcoming challenges in building a reliable CDC system. Speaker: Sigal Shaharabani, Technical Group Leader & Couchbase Ambassador Speaker Profile: Sigal Shaharabani is a Technical Leader and a Group Leader at Tikal and a Couchbase ambassador, with a great passion for backend and data systems. She started her technological career way back in 1996, working for the government, a variety of corporations, and small start-ups. Her favorite programming language is Kotlin, and will love any excuse to talk about it. In her spare time she enjoys swimming and Israeli folk dancing.70Views0likes0Comments