We have a lot of spark sql scripts where the table and field names are not in the correct CASE which is causing these spark notebooks to fail. Can we turn off the Case Sensitivity for spark sql?
4 Comments
- yor_braakmanNew Member
Please adhere to the standard for open source.
For example: “jack” can be used as a verb, meaning to lift or raise something with a jack (e.g., “to jack up a car”).
On the other hand, “Jack” with a capital “J” is typically a proper noun, often used as a name.
These are different sort of words.
If anyone wants to diverge from the standard they can easily add this code:
spark.conf.set('spark.sql.caseSensitive', False)
There's a lot of material out there on collation.
You do not want to force case insensitivity on the rest of the world in 2024.
- fbcideas_migusrNew Member
I find the argument against being case insensitive a bit ludicrous. The general de facto "standard" in SQL is actually to be case insensitive on both table and column names.
For some reason, Apache Spark is inconsistent here and is only case insensitive on column names, at least by default.
The Microsoft world view is more human-friendly and is generally case insensitive across the board. So, Microsoft should definitely be consistent here and be case insensitive by default.
Of course, one can make things more restrictive if they want, but in general default settings should appeal to the majority and not the exception.
- nishalitNew MemberThis is a standard for open source Spark, but the Fabric product group considers the idea worthwhile but it has not been planned yet. The community is encouraged to continue voting and the product team will regularly review these ideas at planning.
- fbcideas_migusrNew MemberStatus added:Needs Votes
Recent ideas
Data Pipelines - Run only selected activities
For debugging and testing pipeline activities during development, allow us to select one or multiple activities and run only the selected pipeline activities. For example, I'm working on editing ...frithjof_v29 minutes agoCommunity ChampionNew606Views11likes2CommentsSemantic model connection bindings should be in source control (Git)
Semantic model data source connection bindings should be source controlled. A semantic model can contain multiple data source references, each of which can be mapped to a separate Fabric data connec...frithjof_v8 hours agoCommunity ChampionNew12Views1like0CommentsBulk changing column names in Visualizations Pane
We often use raw/api column names or measures with a set nomenclature to be consistent and to keep track of them but we do not want to display these names in the visuals. Currently we have to change ...vishal14019711 hours agoFrequent VisitorNew6Views0likes0Comments