Skip to content

RootModel is wrongly shaped by the PySpark codegen #583

Description

Summary

A field typed as a RootModel subclass validates and serializes on the Pydantic side, but the PySpark codegen extracts it as a struct with a root member. Nothing fails at codegen time, the mismatch only surfaces when validating real data.

Pydantic behavior (correct)

A hypothetical per-vehicle-type toll charge on a road segment:

from pydantic import BaseModel, RootModel

class TollChargesByVehicleType(RootModel[dict[str, int]]):
    """Toll charge keyed by vehicle type (e.g. car, hgv, bus)."""

class RoadSegment(BaseModel):
    road_class: str
    toll_charges: TollChargesByVehicleType | None = None

s = RoadSegment.model_validate(
    {"road_class": "motorway", "toll_charges": {"car": 250, "hgv": 900}}
)
s.model_dump() # {'road_class': 'motorway', 'toll_charges': {'car': 250, 'hgv': 900}}

Pydantic validates the bare root value and re-emits it without any wrapper. The overture-schema CLI therefore handle a RootModel-typed field with zero issues.

PySpark Codegen behavior (wrong shape)

The pyspark codegen extraction does not treat RootModel differently. is_model_class is just issubclass(obj, BaseModel), and RootModel is a BaseModel subclass, so extraction produces:

toll_charges: ModelRef(RecordSpec(name='TollChargesByVehicleType',
                fields=[FieldSpec(name='root',
                                  shape=MapOf(key=Primitive('str'), ...))]))

i.e. the root field becomes a real struct member. The generated Spark schema would declare

toll_charges: struct<root: map<string,int>>

while actual Parquet data carries the bare map<string,int>.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions