Summary
A field typed as a RootModel subclass validates and serializes on the Pydantic side, but the PySpark codegen extracts it as a struct with a root member. Nothing fails at codegen time, the mismatch only surfaces when validating real data.
Pydantic behavior (correct)
A hypothetical per-vehicle-type toll charge on a road segment:
from pydantic import BaseModel, RootModel
class TollChargesByVehicleType(RootModel[dict[str, int]]):
"""Toll charge keyed by vehicle type (e.g. car, hgv, bus)."""
class RoadSegment(BaseModel):
road_class: str
toll_charges: TollChargesByVehicleType | None = None
s = RoadSegment.model_validate(
{"road_class": "motorway", "toll_charges": {"car": 250, "hgv": 900}}
)
s.model_dump() # {'road_class': 'motorway', 'toll_charges': {'car': 250, 'hgv': 900}}
Pydantic validates the bare root value and re-emits it without any wrapper. The overture-schema CLI therefore handle a RootModel-typed field with zero issues.
PySpark Codegen behavior (wrong shape)
The pyspark codegen extraction does not treat RootModel differently. is_model_class is just issubclass(obj, BaseModel), and RootModel is a BaseModel subclass, so extraction produces:
toll_charges: ModelRef(RecordSpec(name='TollChargesByVehicleType',
fields=[FieldSpec(name='root',
shape=MapOf(key=Primitive('str'), ...))]))
i.e. the root field becomes a real struct member. The generated Spark schema would declare
toll_charges: struct<root: map<string,int>>
while actual Parquet data carries the bare map<string,int>.
Summary
A field typed as a
RootModelsubclass validates and serializes on the Pydantic side, but the PySpark codegen extracts it as a struct with arootmember. Nothing fails at codegen time, the mismatch only surfaces when validating real data.Pydantic behavior (correct)
A hypothetical per-vehicle-type toll charge on a road segment:
Pydantic validates the bare root value and re-emits it without any wrapper. The
overture-schemaCLI therefore handle a RootModel-typed field with zero issues.PySpark Codegen behavior (wrong shape)
The pyspark codegen extraction does not treat RootModel differently.
is_model_classis justissubclass(obj, BaseModel), andRootModelis aBaseModelsubclass, so extraction produces:i.e. the
rootfield becomes a real struct member. The generated Spark schema would declarewhile actual Parquet data carries the bare
map<string,int>.