Skip to content

FEAT: sparkeks script_uri and entry_point on the SQL wrapper - #143

Merged
Nitin Bhakar (nitinbhakar) merged 7 commits into
mainfrom
feature/sparkeks-pyspark-entrypoint
Sep 8, 2026
Merged

Nitin Bhakar (nitinbhakar) merged 7 commits into
mainfrom
feature/sparkeks-pyspark-entrypoint

Conversation

@nitinbhakar

@nitinbhakar Nitin Bhakar (nitinbhakar) commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Summary

sparkeks can submit a PySpark job on the existing SQL wrapper path. The job names code already on S3; the plugin does not upload query.sql for that case.

  • parameters.script_uri: skip the query.sql upload and pass this object through as query_uri (.py, .zip, .tar.gz, or any other S3 object).
  • parameters.entry_point: on the SQL wrapper, prepended as the next argv after the usual app query_uri user result prefix (path inside an archive). JAR jobs still use entry_point as the main class.
  • Command-level image override so a command can pin a Spark image other than the cluster default.
  • Required spark.sql.extensions are merged after job properties so they cannot be cleared at submit time.

SQL jobs with no script_uri are unchanged: the plugin still uploads query.sql.

Job examples

Same sparkeks command as SQL (Python wrapper). The wrapper inspects query_uri.

1. Single file on S3

POST /api/v1/job
{
  "name": "pyspark-script-job",
  "version": "1.0.0",
  "command_criteria": ["type:pyspark-eks"],
  "cluster_criteria": ["type:spark"],
  "context": {
    "parameters": {
      "script_uri": "s3://example-bucket/pyspark/scripts/job.py"
    }
  }
}

2. Archive on S3

POST /api/v1/job
{
  "name": "pyspark-archive-job",
  "version": "1.0.0",
  "command_criteria": ["type:pyspark-eks"],
  "cluster_criteria": ["type:spark"],
  "context": {
    "parameters": {
      "script_uri": "s3://example-bucket/pyspark/jobs/abc123.zip",
      "entry_point": "src/job.py"
    }
  }
}

Test plan

Submitting a python file

image

Submitting a bundle

image

Testing exsisting Spark SQL logic

image

Verifying the output of pyspark table creation

image
  • go test ./internal/pkg/object/command/sparkeks/...
  • Submit a .py via parameters.script_uri (no query.sql upload)
  • Submit a zip/tar.gz via script_uri + entry_point
  • Confirm SQL jobs with no script_uri still upload query.sql

Enable CI-published zip bundles and ingest-style single-file submits via
parameters.script_uri, with command-level image override and safer Ranger
extension merging.
The command image pin is a generic Spark submit option, not a DS-specific path.
@nitinbhakar
Nitin Bhakar (nitinbhakar) marked this pull request as ready for review September 4, 2026 07:08
Comment thread internal/pkg/object/command/sparkeks/sparkeks.go Outdated
Comment thread internal/pkg/object/command/sparkeks/sparkeks.go Outdated
Comment thread internal/pkg/object/command/sparkeks/entrypoint.go Outdated
Comment thread internal/pkg/object/command/sparkeks/entrypoint.go Outdated
Comment thread internal/pkg/object/command/sparkeks/entrypoint.go Outdated
prasadlohakpure
prasadlohakpure previously approved these changes Sep 8, 2026
@nitinbhakar Nitin Bhakar (nitinbhakar) changed the title FEAT: PySpark bundle and script_uri support in sparkeks FEAT: sparkeks script_uri and entry_point on the SQL wrapper Sep 8, 2026
prasadlohakpure
prasadlohakpure previously approved these changes Sep 8, 2026
@nitinbhakar
Nitin Bhakar (nitinbhakar) merged commit dbd632d into main Sep 8, 2026
7 checks passed
@nitinbhakar
Nitin Bhakar (nitinbhakar) deleted the feature/sparkeks-pyspark-entrypoint branch September 8, 2026 16:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants