[SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #23238

seancxmao · 2018-12-05T15:16:54Z

What changes were proposed in this pull request?

#22148 introduces a behavior change. According to discussion at #22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4.

How was this patch tested?

N/A

…e field resolution when reading from Parquet

SparkQA · 2018-12-05T15:33:02Z

Test build #99732 has finished for PR 23238 at commit 5bbcf41.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

seancxmao · 2018-12-06T09:16:20Z

@HyukjinKwon Would you please kindly take a look at this when you have time?

dongjoon-hyun · 2018-12-07T06:49:23Z

docs/sql-migration-guide-upgrade.md

@@ -141,6 +141,8 @@ displayTitle: Spark SQL Upgrading Guide

  - In Spark version 2.3 and earlier, HAVING without GROUP BY is treated as WHERE. This means, `SELECT 1 FROM range(10) HAVING true` is executed as `SELECT 1 FROM range(10) WHERE true`  and returns 10 rows. This violates SQL standard, and has been fixed in Spark 2.4. Since Spark 2.4, HAVING without GROUP BY is treated as a global aggregate, which means `SELECT 1 FROM range(10) HAVING true` will return only one row. To restore the previous behavior, set `spark.sql.legacy.parser.havingWithoutGroupByAsWhere` to `true`.

+  - In version 2.3 and earlier, when reading from a Parquet data source table, Spark always returns null for any column whose column names in Hive metastore schema and Parquet schema are in different letter cases, no matter whether `spark.sql.caseSensitive` is set to true or false. Since 2.4, when `spark.sql.caseSensitive` is set to false, Spark does case insensitive column name resolution between Hive metastore schema and Parquet schema, so even column names are in different letter cases, Spark returns corresponding column values. An exception is thrown if there is ambiguity, i.e. more than one Parquet column is matched. This change also applies to Parquet Hive tables when `spark.sql.hive.convertMetastoreParquet` is set to true.


Hi, @seancxmao . Maybe, the followings?

- `spark.sql.caseSensitive` is set to true or false + `spark.sql.caseSensitive` is set to `true` or `false`

- `spark.sql.caseSensitive` is set to false + `spark.sql.caseSensitive` is set to `false`

- `spark.sql.hive.convertMetastoreParquet` is set to true + `spark.sql.hive.convertMetastoreParquet` is set to `true`

@dongjoon-hyun Good suggestions. I have fixed them with a new commit.

dongjoon-hyun · 2018-12-07T06:49:54Z

Thank you for adding this to the migration doc. Also, please add [DOC] in the PR title.
cc @gatorsmile .

SparkQA · 2018-12-07T10:08:54Z

Test build #99823 has finished for PR 23238 at commit bf0bb9c.

This patch passes all tests.
This patch merges cleanly.
This patch adds no public classes.

dongjoon-hyun · 2018-12-09T01:52:37Z

Merged to master and branch-2.4.

dongjoon-hyun · 2018-12-09T01:54:30Z

Thanks, @seancxmao .

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? #22148 introduces a behavior change. According to discussion at #22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes #23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <seancxmao@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org> (cherry picked from commit 55276d3) Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>

seancxmao · 2018-12-09T04:05:29Z

Thank you! @dongjoon-hyun

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? apache#22148 introduces a behavior change. According to discussion at apache#22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes apache#23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <seancxmao@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>

…ive field resolution when reading from Parquet ## What changes were proposed in this pull request? apache#22148 introduces a behavior change. According to discussion at apache#22184, this PR updates migration guide when upgrade from Spark 2.3 to 2.4. ## How was this patch tested? N/A Closes apache#23238 from seancxmao/SPARK-25132-doc-2.4. Authored-by: seancxmao <seancxmao@gmail.com> Signed-off-by: Dongjoon Hyun <dongjoon@apache.org> (cherry picked from commit 55276d3) Signed-off-by: Dongjoon Hyun <dongjoon@apache.org>

[SPARK-25132][SQL][FOLLOWUP] Update migration doc for case-insensitiv…

5bbcf41

…e field resolution when reading from Parquet

dongjoon-hyun reviewed Dec 7, 2018

View reviewed changes

seancxmao changed the title ~~[SPARK-25132][SQL][FOLLOWUP] Add migration doc for case-insensitive field resolution when reading from Parquet~~ [SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet Dec 7, 2018

quote boolean values

bf0bb9c

asfgit closed this in 55276d3 Dec 9, 2018

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #23238

[SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #23238

seancxmao commented Dec 5, 2018

SparkQA commented Dec 5, 2018

seancxmao commented Dec 6, 2018

dongjoon-hyun Dec 7, 2018

seancxmao Dec 7, 2018

dongjoon-hyun commented Dec 7, 2018 •

edited

Loading

SparkQA commented Dec 7, 2018

dongjoon-hyun commented Dec 9, 2018

dongjoon-hyun commented Dec 9, 2018

seancxmao commented Dec 9, 2018

		@@ -141,6 +141,8 @@ displayTitle: Spark SQL Upgrading Guide

		- In Spark version 2.3 and earlier, HAVING without GROUP BY is treated as WHERE. This means, `SELECT 1 FROM range(10) HAVING true` is executed as `SELECT 1 FROM range(10) WHERE true` and returns 10 rows. This violates SQL standard, and has been fixed in Spark 2.4. Since Spark 2.4, HAVING without GROUP BY is treated as a global aggregate, which means `SELECT 1 FROM range(10) HAVING true` will return only one row. To restore the previous behavior, set `spark.sql.legacy.parser.havingWithoutGroupByAsWhere` to `true`.

		- In version 2.3 and earlier, when reading from a Parquet data source table, Spark always returns null for any column whose column names in Hive metastore schema and Parquet schema are in different letter cases, no matter whether `spark.sql.caseSensitive` is set to true or false. Since 2.4, when `spark.sql.caseSensitive` is set to false, Spark does case insensitive column name resolution between Hive metastore schema and Parquet schema, so even column names are in different letter cases, Spark returns corresponding column values. An exception is thrown if there is ambiguity, i.e. more than one Parquet column is matched. This change also applies to Parquet Hive tables when `spark.sql.hive.convertMetastoreParquet` is set to true.

[SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #23238

[SPARK-25132][SQL][FOLLOWUP][DOC] Add migration doc for case-insensitive field resolution when reading from Parquet #23238

Conversation

seancxmao commented Dec 5, 2018

What changes were proposed in this pull request?

How was this patch tested?

SparkQA commented Dec 5, 2018

seancxmao commented Dec 6, 2018

dongjoon-hyun Dec 7, 2018

Choose a reason for hiding this comment

seancxmao Dec 7, 2018

Choose a reason for hiding this comment

dongjoon-hyun commented Dec 7, 2018 • edited Loading

SparkQA commented Dec 7, 2018

dongjoon-hyun commented Dec 9, 2018

dongjoon-hyun commented Dec 9, 2018

seancxmao commented Dec 9, 2018

dongjoon-hyun commented Dec 7, 2018 •

edited

Loading