Forum Discussion
dataflow source change from aws redshift to databricks
- 8 months ago
Hi bhavyamalik1,
You can change the data source in your dataflow, but it requires recreating the connection. Here's how you can try to approach it:
Option 1: Edit directly in Power BI Service (Recommended)
- Go to your workspace and open the dataflow for editing
- In Power Query Editor, select the query connected to Redshift
- Click Data source settings or Configure connection
- Unfortunately, you cannot switch connector types (Redshift → Databricks) directly
- You'll need to delete the current source step and add a new Databricks connection
- Right-click the Source step → Delete
- Get Data → Databricks → enter your connection details
- Reapply your transformation steps (they should still be there)
Option 2: Use the JSON file (Advanced)
- Open the downloaded JSON file
- Find the data source section (look for "Amazon Redshift" references)
- Replace with Databricks connection string format:
- Change the connector from AmazonRedshift to Databricks
- Update server, database, and authentication details
- This is risky - one syntax error breaks the dataflow
Option 3: Recreate the dataflow (Safest)
- Create a new dataflow
- Connect to Databricks
- Copy/paste the M code from your existing queries (from the JSON or Advanced Editor)
- Update reports to point to the new dataflow
Best regards!
PS: If you find this post helpful consider leaving kudos or mark it as solution
Hi bhavyamalik1
I hope you are doing well!!
You don’t need to modify the exported JSON file. Dataflow definitions aren’t designed to be edited directly, and manual changes may break the dataflow.
The correct way is to update the connection through Power Query Online:
1. Open the Power BI Service
2. Navigate to the workspace that contains the dataflow
3. Select Edit dataflow
4. In Power Query Online, open each query that currently uses Amazon Redshift
5. Update the Source step to use the Azure Databricks connector instead
6. Provide the Databricks connection details (server, HTTP path, authentication)
7. Save the dataflow and validate the refresh
If the schema between Redshift and Databricks is different, it’s often cleaner to:
Create a new query using Databricks as the source
Reapply the existing transformation steps
Replace the old query once validated
Editing the JSON file is not supported and is not recommended.
This approach keeps the dataflow stable, supported, and easy to maintain.
If this help you mark as "Solution" and put kudo to help other