field_path (added in #518, now #569) models a path to a value through a nested schema. Today it splits paths into three classes — ScalarPath (no iteration), ArrayPath (array iteration), MapPath (one map projection, struct-only leaf) — and forbids anything that mixes or nests containers. Victor Schappert (@vcschapp) flagged in review that those limits feel arbitrary, and they are: the codegen generates validation for the schema's full type space, not for the shape the schema happens to have today, and these limits bake the current shape in as a ceiling. The direction is to generalize, not to trim field_path down to what's currently used.
What's restricted today
A MapPath requires the map to be reachable without array iteration and its leaf to be struct-only. parse rejects dict[K, list[V]] (an array marker after a map projection) and multiple map markers; promote_terminal_map / promote_terminal_array raise NotImplementedError for a map under an array element, a map under a map element, or an array under a map element. In both Pydantic and Spark those shapes are ordinary and nest arbitrarily deep, so nothing about the data forces the restriction — it comes from the path grammar and the three-class split.
Why it hasn't surfaced yet
The only map in the schema right now is CommonNames (dict[LanguageTag, StrippedString]), a dict[K, scalar] that MapPath handles. So the guards above are unexercised, not validated as acceptable — the first richer map field (a dict[K, list], a dict[K, Model] reached through an array, a nested map) hits a NotImplementedError at codegen time. Generalizing now keeps that from being a surprise later.
The design is already half-generalized
MapPath carries a non-empty-leaf case for dict[K, Model] (e.g. subs{value}.label) that no field exercises yet, while the container axis stays stubbed with NotImplementedError. The leaf direction was generalized and the nesting direction wasn't; the fix is to finish the generalization, not retreat to dict[K, scalar].
The work
- Collapse the taxonomy to
Direct (no iterating segment — today's ScalarPath) versus a single iterated path type whose segments freely mix ArraySegment and MapSegment. ArrayPath and MapPath merge, and their helper properties (array_chunks, iter_struct_paths, map_column, leaf) generalize to walk a mixed segment sequence.
- Extend the string grammar to encode the shapes
parse rejects today: dict[K, list], nested maps, and maps under array elements. Drop the corresponding ValueErrors and the promote_* NotImplementedErrors.
- The
ScalarPath naming point from review folds in here. Once the split is Direct versus Iterated, there's no per-container member name left to get wrong — ScalarPath named its terminal (a scalar, though it also names struct prefixes via column_prefix) while ArrayPath / MapPath named the fan-out, a mixed axis the collapse removes.
Not a blocker for #569.
field_path(added in #518, now #569) models a path to a value through a nested schema. Today it splits paths into three classes —ScalarPath(no iteration),ArrayPath(array iteration),MapPath(one map projection, struct-only leaf) — and forbids anything that mixes or nests containers. Victor Schappert (@vcschapp) flagged in review that those limits feel arbitrary, and they are: the codegen generates validation for the schema's full type space, not for the shape the schema happens to have today, and these limits bake the current shape in as a ceiling. The direction is to generalize, not to trimfield_pathdown to what's currently used.What's restricted today
A
MapPathrequires the map to be reachable without array iteration and its leaf to be struct-only.parserejectsdict[K, list[V]](an array marker after a map projection) and multiple map markers;promote_terminal_map/promote_terminal_arrayraiseNotImplementedErrorfor a map under an array element, a map under a map element, or an array under a map element. In both Pydantic and Spark those shapes are ordinary and nest arbitrarily deep, so nothing about the data forces the restriction — it comes from the path grammar and the three-class split.Why it hasn't surfaced yet
The only map in the schema right now is
CommonNames(dict[LanguageTag, StrippedString]), adict[K, scalar]thatMapPathhandles. So the guards above are unexercised, not validated as acceptable — the first richer map field (adict[K, list], adict[K, Model]reached through an array, a nested map) hits aNotImplementedErrorat codegen time. Generalizing now keeps that from being a surprise later.The design is already half-generalized
MapPathcarries a non-empty-leaf case fordict[K, Model](e.g.subs{value}.label) that no field exercises yet, while the container axis stays stubbed withNotImplementedError. The leaf direction was generalized and the nesting direction wasn't; the fix is to finish the generalization, not retreat todict[K, scalar].The work
Direct(no iterating segment — today'sScalarPath) versus a single iterated path type whosesegmentsfreely mixArraySegmentandMapSegment.ArrayPathandMapPathmerge, and their helper properties (array_chunks,iter_struct_paths,map_column,leaf) generalize to walk a mixed segment sequence.parserejects today:dict[K, list], nested maps, and maps under array elements. Drop the correspondingValueErrors and thepromote_*NotImplementedErrors.ScalarPathnaming point from review folds in here. Once the split isDirectversusIterated, there's no per-container member name left to get wrong —ScalarPathnamed its terminal (a scalar, though it also names struct prefixes viacolumn_prefix) whileArrayPath/MapPathnamed the fan-out, a mixed axis the collapse removes.Not a blocker for #569.