Skip to content

Wrap nullable struct fields in an Avro union in to_avro() - #283

Open
Farid841 wants to merge 1 commit into
astrolabsoftware:mainfrom
Farid841:fix/nullable-struct-avro-schema
Open

Farid841 wants to merge 1 commit into
astrolabsoftware:mainfrom
Farid841:fix/nullable-struct-avro-schema

Conversation

@Farid841

@Farid841 Farid841 commented Sep 8, 2026

Copy link
Copy Markdown

_parse_struct() wraps nullable array and map fields in a [type, null] Avro union via _is_nullable(), but nullable struct fields (e.g. the 'candidate' field of a ZTF alert) were passed through as plain records. Spark's own to_avro() serializer always writes a union discriminator byte for nullable fields regardless of the published schema, so any strict Avro reader following the schema this function produces fails to decode nullable struct fields with an out-of-range/index error.

I hit this while building a Kubernetes AI inference feature for the
Fink science portal: a preprocessing container consumes ZTF alerts straight
from a topic whose schema was published with to_avro(), and fails to
deserialize every message because of the candidate field specifically:

"error": "list index out of range", "topic": "ftransfer_ztf_...",
"partition": 6, "offset": 0

_parse_struct() wraps nullable array and map fields in a [type, null]
Avro union via _is_nullable(), but nullable struct fields (e.g. the
'candidate' field of a ZTF alert) were passed through as plain records.
Spark's own to_avro() serializer always writes a union discriminator
byte for nullable fields regardless of the published schema, so any
strict Avro reader following the schema this function produces fails
to decode nullable struct fields with an out-of-range/index error.

Downstream, this has been worked around ad hoc by patching the
generated schema JSON after the fact (see
astrolabsoftware/ztf.fink-portal.org's spark_ztf_inference_feed.py);
fixing it here removes the need for that per-caller patch.
Copilot AI lite review requested due to automatic review settings September 8, 2026 07:05

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The change is small, consistent with existing nullable handling for arrays/maps, and directly addresses the described Avro decoding failure mode.

Pull request overview

This PR fixes Avro schema generation for Spark to_avro() compatibility by ensuring nullable struct fields are encoded as Avro unions, matching Spark’s serialization behavior and preventing strict Avro readers from failing on nullable records (e.g., ZTF candidate).

Changes:

  • Wrap nullable struct fields produced by _parse_struct() using the existing _is_nullable() union logic (consistent with array/map handling).
File summaries
File Description
fink_utils/spark/schema_converter.py Ensures nullable nested structs are emitted as [recordType, "null"] unions so readers expecting unions can decode Spark-serialized data.
Review details
  • Files reviewed: 1/1 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants