Provide the option of not removing duplicates automatically when creating R visualizations or provide the ability to create R datasets using the same syntax as shown in the comments when creating an R visualization
17 Comments
- gdeckler4New MemberThe comments of an R visualization show: #dataset <- data.frame(Column) However, I cannot use the same syntax to create my own data frame.
- boefraty
Microsoft Employee
The workaround here is to add "ID column" to the data (don't use it in R script) - efglynn
Advocate IV
I shouldn't have to add an ID or key field to get all the data. If I want to remove duplicates in R, it's trivial with the "unique" statement. - efglynn
Advocate IV
When linking to an SQL Server Analysis Services database cube, I don't directly have access to the right keys that make records unique to block duplicates from being removed. Therefore, when using cubes it may not be possible to get accurate data in R in some cases. Many statistics/visualizations are worthless when duplicates have been removed. Can someone explain why removing duplicates was ever a good idea? - jo_varneyNew MemberI want to use R to create a histogram. I add one column, and then it removes all the duplicates, which provides a completely inaccurate histogram. This is pretty stilly - and potentially problematic if someone uses this without noticing. Yes, I can add extra columns, or create an ID column, but I don't want to. I don't want the program to remove duplicates, just because it sees fit. There are times when it isn't appropriate - and as the analyst I want that choice. Also, I want to write the simplest code, and that should involve only one column for a histogram.
- george_fullegarNew MemberI have no idea why this feature isn't standard behaviour - it's trivial in R to remove duplicates from a dataset if that behaviour is desired
- MAwbre
Advocate I
Why does it remove duplicates by default? When performing univariate qualitative analysis, I want to be able to drop in a single qualitative field. This means I WANT duplicates and having keys complicates the analysis meaninglessly. That automatic removal should be made an option. I believe that it was added because of the limitation of R scripts in Power BI to 150k rows. The removal of duplicates by Power BI (in what I assume to be some sort of pre-processor directive like call judging by the invalid code syntax shows in the editor) probably helps mitigate that limitation in certain types of data sets. Unfortunately, without the option to turn off that "pre-processor" like call, an entire segment of potential analysis is complicated or even impossible (if the original data set has no key). - aaron_s_creightNew MemberThe work around for this is the use of an Index column. However, this does not work if the data is coming from different tables!!! I would have to create a new table with a index column for that new table containing the variables of interest. This would defeat the purpose of using a rational data structure and other aspects.
- sleptwithyomamaNew MemberMicrosoft always adds useless features that always are dumb. The very least, for this stupid feature, is to have it disable-possible.
- renvaldas1New MemberPlease do the same for Python!
Recent ideas
Allow admins to choose the preferred connection type for automatic binding
Currently, when the same data source qualifies for both a Cloud Connection/Personal Cloud Connection (PCC) and an Enterprise gateway connection, Power BI may identify multiple eligible connection can...v-viputta13 hours agoMicrosoft EmployeeNew187Views57likes0CommentsSupport Dynamic Base URL Parameterization in Fabric Web/HTTP Connections
a { text-decoration: none; color: #464feb; } tr th, tr td { border: 1px solid #e6e6e6; } tr th { background-color: #f5f5f5; } Currently, Microsoft Fabric Web/HTTP connections require a static base U...pandegautami4 hours agoRegular VisitorNew8Views1like0Comments