Forum Discussion
More rows than filtered in Power query
- 6 years ago
I think the problem is the ODBC driver. I just connected to my SQL server via ODBC instead of directly and it does not support query folding in Power Query. So what that means is Power Query is importing your ENTIRE TABLE contents for each table you connect to, and then it works with everything in RAM. It is very inefficient, and little better than working with CSV files. In fact, it is probably about the same, except your tables are coming in structured and the ODBC driver is hopefully passing along data types and other relevant metadata.
But Power Query isn't generating nice neat SQL statements and sending it back for processing. That is assuming it is a relational database to begin with. It may not be, in which case no folding will happen regardless of connection type.
You can validate this though. Pull in 1 table. Do 1 simple filter. Right-click on that step. Does it say "View Native Query?" If that is grayed out, you aren't getting the advantage of a relational database connection. If it is, my whole theory on your speed issue is shot - other than I don't know what your ODBC driver is doing exactly and it could be doing a lot of interpretation and processing slowing things down.
I had an issue a few years ago I had to connect to an old SQL server (2000 I believe) that didn't support features required by Power Query so I had to pull in entire tables for Power Query to work with. It was on site, so processing about 4M records across about 10 tables and doing joins and what-not took 20-25min start to finish on a PC with 16GB of RAM to return 7,000 relevant records. Your record count is smaller, but I've no clue how many columns you have or the column content. Plus you said you were working remotely, so the transmission speeds wouldn't match what I was getting via ethernet.
If it is a relational database on the other end, and you cannot do a direct connection vs bypassing the ODBC driver, you could create a view on the server. You can see here what data sources Power BI supports natively (no need for ODBC). Sources like Amazon Redshift, IBM DB2, Oracle, SQL, SAP HANA, PostgreSQL, Vertica and many others support query folding. A view would let you apply filters and remove unnecessary columns, radically reducing data transmission times and local processing times.
You don't need to know how to code to do joins/merges in Power Query. It is all in the UI. See this article for more. As for the other issues, I'm having trouble understanding what you are trying to accomplish with your description, so perhaps sharing your file would help.
Hi,
Can the Indexes in Query Editor solve the issue? meaning when I refresh in desktop the uploading query runs through only filtered rows instead of all the 100 rows in my table. so the refresh does not take longer!