"shredding" is copying fields out of the json objects into their own columns so you can query them fast without parsing every json object
the first thing they tried was "static shredding" where they do this for a hardcoded list of fields
the new nice thing is "dynamic shredding" where they pick which fields to copy out based on observed usage patterns, at the time they build the parquet files.
Is "shredding" an industry accepted term? We've been doing this for years at my current job and we call it "exploding". Same idea though: nested JSON objects get dunder path separators and arrays get a new table with a foreign key join back to the parent row.
I'm not sure in what context 32 or 128 are large numbers of keys. We explode out tens of thousands of columns dynamically. The only exceptions we have to make are for fields when the cardinality is as high as the row count, e.g., when there's a UUID or something in the actual JSON path. After five years of exploding we're only starting to hit this enough now where I'm wondering whether we should implement a detection and pivot around such sprawl to turn them into values. If anyone's familiar with considering this please point me to anything you've found/discovered regarding it.
0x2ba22e11 | 13 hours ago
Nice! Condensed:
peter-leonov | 5 hours ago
Thanks for sharing! Did you/they compare this outside-the-format approach with an in-the-format extension? Something like what CH does for JSON.
ryan-duve | 49 minutes ago
Is "shredding" an industry accepted term? We've been doing this for years at my current job and we call it "exploding". Same idea though: nested JSON objects get dunder path separators and arrays get a new table with a foreign key join back to the parent row.
I'm not sure in what context 32 or 128 are large numbers of keys. We explode out tens of thousands of columns dynamically. The only exceptions we have to make are for fields when the cardinality is as high as the row count, e.g., when there's a UUID or something in the actual JSON path. After five years of exploding we're only starting to hit this enough now where I'm wondering whether we should implement a detection and pivot around such sprawl to turn them into values. If anyone's familiar with considering this please point me to anything you've found/discovered regarding it.