apache iceberg
6 TopicsBring your Snowflake Iceberg tables to OneLake with Fabric mirroring
Microsoft OneLake is the single, unified, logical data lake for your whole organization. If you already have Apache Iceberg tables created in Snowflake and stored in Amazon S3, Google Cloud Storage, or Azure Storage, you don't need to copy them or stand up a pipeline to use them across Fabric. With Fabric mirroring, you can bring your Snowflake Iceberg tables into OneLake and start querying them in minutes, all with no data duplication. Here's how it works and how to get started. How are Iceberg tables mirrored to OneLake? When you mirror a Snowflake database into OneLake, your Iceberg tables are handled differently from the rest of the tables being mirrored. Instead of copying the data, mirroring replicates each Iceberg table's metadata into OneLake using a shortcut that points back to the storage holding your Iceberg data, so your Parquet data files stay exactly where they are. Today, this works for Snowflake Iceberg tables stored in your own Amazon S3, Google Cloud Storage, or Azure Storage location — the storage you connect Fabric to during mirroring setup. Support for Iceberg tables that use Snowflake-managed storage is coming soon. OneLake then automatically virtualizes those Iceberg tables as Delta Lake, so every Fabric engine (SQL, Spark, Power BI, and more) can read them natively. You get one managed experience that keeps your tables in sync, instead of creating and maintaining a shortcut for every table by hand. Before you start A few prerequisites make setup smooth: Snowflake permissions. The user configuring mirroring needs CREATE STREAM, SELECT, SHOW tables, and DESCRIBE tables on the source database, plus a role with access to the Snowflake instance. Networking.If your Snowflake instance is behind a firewall or in a private network, set up a virtual network data gateway or an on-premises data gateway so Fabric can reach it. Fabric capacity. You'll need an existing Fabric capacity, or you can start a trial. How to get started Bringing your Iceberg tables into OneLake takes just a few steps: In your Fabric workspace, open the Create hub and select the Mirrored Snowflake card. Name your mirrored database and connect to your Snowflake instance using username/password or Microsoft Entra single sign-on. Choose the tables you want to mirror: mirror everything or pick individual tables. For any Iceberg tables, connect Fabric to the underlying Amazon S3, Google Cloud Storage, or Azure Storage location where they're stored. A single storage connection is used, so make sure the Iceberg tables you select are reachable through the same location. Select Create. Fabric provisions a mirrored database item along with an auto-generated, read-only SQL analytics endpoint. That's it. Your Iceberg tables now appear as Delta Lake tables in OneLake, ready to use across Fabric. For the full detailed walkthrough, see the Snowflake mirroring tutorial. Good to know A few details worth keeping in mind: Table types. Mirroring supports Snowflake Iceberg tables alongside standard native tables. External, transient, temporary, and dynamic tables aren't supported, and you can mirror up to 1,000 tables per database. Security. Granular permissions from Snowflake aren't carried over — re-apply any object-level security in Fabric. Cost. Fabric compute for mirroring is free, and mirroring storage is free up to a capacity-based limit. Query compute (SQL, Power BI, Spark) is billed at normal rates, and Snowflake-side compute applies when data changes are read. For the complete list, explore the Snowflake mirroring limitations. What you unlock in OneLake Once your tables are in OneLake, the whole Fabric platform is open to them — all on data that never left its original location. Query with T-SQL through the SQL analytics endpoint, build Spark notebooks, power reports with Direct Lake in Power BI, run data science workloads, and share data across your organization. Next steps Give it a try on your own Snowflake Iceberg data and let us know what you think. Dive into the documentation, share feedback on Fabric Ideas, and join the conversation in the Fabric Community.1.1KViews0likes0CommentsMicrosoft OneLake and Snowflake interoperability (Generally Available)
Data teams today are under extraordinary pressure. Expectations around analytics and AI have never been higher, yet enterprise data continues to live across a patchwork of systems, tools, and platforms. The result is friction, duplication, and complexity, making it harder for data teams to provide a unified, real-time view of their business. Microsoft and Snowflake committed to solving this challenge with a shared vision: empower our customers to access, analyze, and share data across platforms without duplication and without locking them into proprietary formats. Today, that vision becomes a reality. At Snowflake’s BUILD conference in London today, we are co-announcing Microsoft OneLake’s interoperability with Snowflake is now generally available, marking a major milestone for thousands of joint customers building modern, open and interoperable data foundations. Adding to the list of already generally available integrations, we are now making the ability to natively store Snowflake-managed Iceberg tables in OneLake, the automatic conversion of Fabric data into Iceberg format for direct access from Snowflake, and new UI experiences in both platforms into general availability. If you’re interested in learning more about our collaboration with Snowflake, I recently sat down with Christian Kleinerman, Executive VP of Product at Snowflake, to discuss our integration and talk broadly about the future of our platforms and the industry: Unlocking full cross-platform interoperability AI has accelerated the need for organizations to unify their data estates, but doing so with traditional, copy-heavy approaches is slow, expensive, and brittle. The now generally available (GA) integration between OneLake and Snowflake gives customers something fundamentally different: the freedom to choose the right architecture, storage, and analytical engine for every scenario, without introducing silos or operational overhead. So, what does this GA mean for you? The ability to bidirectionally read Iceberg data managed by Snowflake or Microsoft Fabric. The ability to natively store Snowflake-managed iceberg tables in Microsoft OneLake. Taken together, you now have full flexibility to store your data in either platform while picking the right tool in either Fabric or Snowflake for any project. Since it’s a single copy of data, any changes you make to your data in one platform are reflected in both. We’re also making new UI elements in both platforms generally available starting next week, including a new Snowflake item in OneLake. This item allows you to access your Snowflake data in OneLake faster without complicated configurations. Snowflake has also introduced new UI that allows users to seamlessly push managed Iceberg tables directly into Fabric, making them instantly discoverable as OneLake items. Check out a quick demo of this new UI in action: Looking ahead: empowering customers on their terms While we are thrilled at the progress we’ve made already, this isn’t the end of our collaboration. We will continue to release new capabilities that help bring our platforms together, like our recently released OneLake table APIs which integrate directly with Snowflake’s catalog-linked database feature. To see the latest in person, I’d encourage you to join us for FabCon & SQLCon 2026 in Atlanta, Georgia. The event takes place from March 16–20, 2026 and will give you a unique opportunity to master the latest innovations and get practical guidance from both Fabric and Snowflake experts. You can register for either FabCon or SQLCon conferences and enjoy full access to both. You can also learn more about our current capabilities through these resources: Watch the keynote broadcast of Snowflake’s BUILD conference in London. Try the Snowflake Quickstart for Iceberg in OneLake and watch the walkthrough video guide for hands on demos. Read the Microsoft Learn documentation on using Snowflake with Iceberg tables in OneLake Explore our OneLake developer guidance for endpoints and SDKs.141KViews1like0CommentsMirroring: SQL Server, PostgresSQL, Cosmos DB, Updates to Snowflake Mirroring and the Mirroring Ecosystem (Generally Available)
Data silos slow innovation. Mirroring in Microsoft Fabric eliminates those barriers by bringing operational data into OneLake without complex ETL pipelines. The result, near real-time AI and BI insights, unified governance, and a single source of truth across your organization – all at no cost! What’s Generally Available and why it matters: SQL Server 2016-2022 & SQL Server 2025 (Generally Available) Native mirroring for SQL Server (Generally Available), including versions 2016–2022 and SQL Server 2025. This capability securely replicates operational data into OneLake in an analytics-ready format, converting it to Delta tables for downstream scenarios like BI, data science, and engineering. Each mirrored database automatically provides a SQL analytics endpoint, enabling familiar T-SQL queries on read-only copies without ETL overhead. Supported environments include on-premises, Azure VMs, and non-Azure clouds, with secure connectivity via on-premises data gateway. To learn more about Mirroring SQL Server, refer to the documentation our check out the in-depth blog post. "With Fabric Mirroring in SQL Server 2025, ExponentHR can effortlessly mirror numerous datasets to fabric, enabling near real-time analytics. This technology has alleviated the need for expensive and complex ETL operations and enables more productivity for our customers. Thanks to SQL Server 2025’s built-in cloud connectivity, we can directly process large amounts of data efficiently and overcome traditional bottlenecks." Brent Carlson, IT Manager, ExponentHR Snowflake Mirroring for Iceberg Tables (Generally Available) Snowflake Mirroring now supports for Apache Iceberg tables in addition to managed tables enabling organizations to bring external Iceberg datasets into OneLake without heavy overhead. This enhancement to Snowflake Mirroring uses shortcut-based mirroring to provide read-through access to Iceberg tables hosted across diverse storage systems such as ADLS Gen2, AWS S3, Google Cloud Storage, and S3-compatible services. You can now view your mirrored managed tables and Iceberg tables in one place and unlock high-performance and AI-ready analytics, open format interoperability, and cost efficiency. To learn more, and a free trial, refer to the Tutorial: Configure Microsoft Fabric mirrored databases from Snowflake documentation. Azure PostgreSQL (Generally Available) Mirroring for Azure Database for PostgreSQL Flexible Server (Generally Available), enables continuous replication of transactional data into OneLake without manual ETL. Mirrored PostgreSQL databases are converted into Delta tables, making them analytics-ready for BI, machine learning, and data engineering workloads. This feature ensures near real-time synchronization, supports schema evolution, and provides a SQL analytics endpoint for querying mirrored data using familiar T-SQL syntax. It’s ideal for scenarios like customer analytics, inventory optimization, and real-time reporting. To learn more, refer to the Mirroring Azure Database for PostgreSQL flexible server documentation. Azure Cosmos DB (Generally Available) Cosmos DB Mirroring (Generally Available) brings globally distributed NoSQL data into OneLake for unified analytics. This GA release supports continuous change capture, converting mirrored containers into Delta tables for downstream analytics while maintaining Cosmos DB’s multi-model flexibility. Key enterprise features include automatic schema inference and support for nested JSON. Use cases span real-time personalization, fraud detection, and IoT telemetry analysis at scale. To learn more, refer to the Mirroring Azure Cosmos DB documentation or check out our in-depth blog post. Qlik joins the Mirroring Partner Ecosystem Qlik's Open Mirroring integration covers over 40 different sources including SAP and DB2 just to name a few. This integration automates and simplifies the extraction and streaming of operational data directly into OneLake. The result is a unified data estate and near real-time data freshness that eliminates complex ETL pipelines, minimizes source system impact, and delivers instant, reliable insights for analytics, data science, and AI use cases, all while reducing manual configuration and maintenance overhead. Eastman has already seen significant efficiencies with mirroring from their SAP Application DB on Oracle systems. “Qlik Replicate's Open Mirroring integration has dramatically reduced end-to-end latency from over 10 minutes to under one minute and moved us from daily operational interventions to sustained, long-term stability. This combination has transformed our ability to deliver real-time insights to Microsoft Fabric targets with the reliability and performance our teams require.” Ben Hobbs, Senior Data Architecture Administrator at Eastman Why this changes the game Accelerated insights with zero ETL. Simplified governance under OneLake. Reduced cost and complexity across hybrid and multi-cloud environments What’s Next Expect more connectors, performance optimizations, and deeper integrations in the months ahead. Stay tuned for roadmap updates and private previews. Engage with us on Fabric Ideas, we love feedback! Try Mirroring in Fabric, sign up for a free trial and get started today!99KViews0likes0CommentsOneLake Table APIs (Preview)
Microsoft OneLake is the unified data lake for your entire organization, built into Microsoft Fabric. It provides a single, open, and secure foundation for all your analytics workloads – eliminating data silos and simplifying data management across domains. The preview of Microsoft OneLake Table APIs, a new way to programmatically manage and interact with your data tables in OneLake! These APIs open the door for developers and data engineers to integrate OneLake seamlessly into their workflows, enabling powerful automation and interoperability with open table formats. OneLake table APIs OneLake is designed to be your organization’s single, unified data lake, and with the new Table API endpoint, you can seamlessly integrate your data into modern, open analytics workflows—making connectivity and management easier than ever. List and fetch details of tables and schemas in OneLake programmatically. Integrate with open-source ecosystems using familiar APIs. Build custom applications and services that interact directly with your OneLake tables. Starting with Apache Iceberg REST Catalog (IRC) In this initial release, OneLake Table APIs support the Iceberg REST Catalog (IRC) specification, giving you a familiar and standards-based way to manage Iceberg tables. If you’re already using Iceberg, you can now connect your tools and applications to OneLake without changing your existing workflows. Soon, we’ll expand support to include Delta Lake operations, which will be covered in an upcoming update – so stay tuned! Getting started Getting started is simple – and the documentation makes it even easier by providing examples for different tools, services, and libraries. You’ll find step-by-step guidance to configure your Iceberg REST Catalog clients to work with OneLake Table APIs – whether you’re using Snowflake, DuckDB, PyIceberg, or other Iceberg-compatible clients. Visit the documentation for OneLake table APIs for Iceberg. Explore examples to see how to: Set up your Iceberg REST Catalog client. Connect to OneLake using the new APIs. Perform common operations like creating and listing tables. Use these examples as inspiration to integrate OneLake into your workflows quickly and confidently. Why it's important By supporting open standards like Apache Iceberg, OneLake continues to deliver on its promise of openness and flexibility. Whether you’re building analytics pipelines, machine learning workflows, or custom data applications, these APIs give you the tools to integrate OneLake into your ecosystem with ease. What’s next? This is just the beginning. Soon, we’ll introduce Delta Lake support to our OneLake table APIs, enabling you to manage Delta tables with the same simplicity and openness. We’re also working on expanding API capabilities for additional operations, opening the door for more integration opportunities. There’s much more to come as we grow the set of APIs – so stay tuned! Have feedback? Your input helps shape the future of OneLake. Try out the new table APIs and let us know what works well and what you’d like to see next. Share your feedback through Microsoft Fabric Ideas. Ready to dive in, refer to the Use Iceberg tables with OneLake documentation and start exploring the possibilities today!87KViews0likes0CommentsHow Microsoft OneLake seamlessly provides Apache Iceberg support for all Fabric Data
Co-authored with Kevin Liu, Apache Iceberg Committer and Principal Engineer at Microsoft. Microsoft Fabric is a unified SaaS data and analytics platform designed for the era of AI. All workloads in Microsoft Fabric use Delta Lake as the standard, open-source table format. With Microsoft OneLake, Fabric’s unified SaaS data lake, customers can unify their data estate across multiple cloud and on-prem systems. We recently announced that OneLake can now transparently serve Delta Lake tables as Apache Iceberg tables. This post covers the details of how we built this feature and enables seamless interoperability across engines and ecosystems. Try out this feature today! Get started: 1. Create or identify a Delta Lake table in OneLake. 2. Use any Iceberg-compatible engine (e.g., Spark, Trino, Snowflake) to query the table. 3. OneLake will automatically serve the table in Iceberg format – no configuration needed. Why This Matters Organizations often rely on diverse engines and tools that favor different open table formats. Delta Lake is a popular choice for Spark-based pipelines and Microsoft Fabric engines, while Iceberg is increasingly adopted by query engines like Trino, Dremio, and Snowflake. With OneLake, there’s no need to choose one format over the other. OneLake empowers openness and interoperability by automatically translating metadata between Delta and Iceberg formats. This means your tables are seamlessly available in both ecosystems—without duplication or compromise. Whether your team prefers Delta or Iceberg, OneLake ensures compatibility, flexibility, and freedom of choice. By enabling on-the-fly conversion from Delta to Iceberg, OneLake eliminates the need for data duplication or format-specific pipelines. You can now: – Read Delta tables as Iceberg without rewriting or copying data. – Use Iceberg-native tools on top of your existing Delta Lake datasets. – Simplify governance and access control with a single source of truth. Example Here we have a table in Fabric Lakehouse stored in the Delta format. Table in Fabric Lakehouse Select ‘View files’ on table to see the table’s metadata. View files belonging to a table in Fabric Lakehouse Both _delta_log and metadata directories should appear. Metadata folders for a table in Fabric Lakehouse Select the metadata directory to view the Iceberg files. Iceberg metadata for table in Lakehouse Navigate to Snowflake and create an external volume pointing to the exampled Lakehouse. Create external volume in Snowflake Next, we create an Iceberg catalog. Create Iceberg catalog in Snowflake Create a table that uses the Iceberg catalog and external volume, specifying the Iceberg metadata file path in OneLake. Create table in Snowflake That's it - now you can query the table! Query table in Snowflake How It Works OneLake uses a feature called table format virtualization to surface tables as both Delta and Iceberg formats. When an Iceberg-compatible engine reads from a OneLake table that is natively in the Delta format, OneLake dynamically generates the necessary Iceberg metadata files. This allows the engine to interpret the table as if it were natively an Iceberg table. Behind the scenes, this feature utilizes Apache XTable for table format metadata conversion. XTable provides cross-table omni-directional interop between the open table formats. We have also enhanced XTable functionality - for example, by converting Delta deletion vectors into Iceberg positional delete files. We look forward to contributing these features upstream to the open-source community. Virtualized Iceberg metadata for a table in Fabric The original Delta table resides in OneLake or an external location (e.g., ADLS, S3). When queried via an Iceberg engine, OneLake serves Iceberg metadata derived from the metadata directory. Here’s the workflow for a sample request. Reading Iceberg metadata for a table in Fabric The source metadata remains in its original Delta structure (in the _delta_log directory). OneLake’s virtualization layer intercepts read requests on the metadata directory and determines that Iceberg-compatible metadata is needed. It then generates Iceberg-compliant metadata on demand, enabling engines to interact with the table as if it were natively Iceberg, without any physical data movement. This approach is designed to be transparent to the table owner and does not modify the original storage location. It maintains consistency with the source table and scales efficiently by producing metadata on demand, when requested by a client. Future Work We’re actively expanding support for additional data types and advanced features across both Delta Lake and Iceberg formats. A key focus is ensuring compatibility with the upcoming Iceberg V3 specification. These enhancements will further improve OneLake’s cross-format interoperability and simplify analytics across engines. Learn More New in OneLake: Access your Delta Lake tables as Apache Iceberg automatically Store and use your Iceberg data in OneLake Microsoft Fabric documentation90KViews0likes0Comments