CGinS — CUDA Ghost in the Shell explores how language models can translate PyTorch operations into CUDA kernels using runtime context and correctness feedback. This repository brings together the research implementation, generated-kernel workspace, and a Jac interface.
Implementation · Research paper · Original documentation · Jac interface guide
flowchart LR
A["Profile<br/>PyTorch operation"] --> B["Capture<br/>tensor context"]
B --> C["Generate<br/>CUDA"]
C --> D["Compile<br/>and validate"]
D -->|Feedback| C
D --> E["Inspect<br/>the kernel"]
The interesting boundary is between generated code and executable evidence: tensor inputs, compilation results, and comparison with the reference operation all participate in the feedback loop.
| Location | Purpose |
|---|---|
| CGinS-gui/src | Generation and optimization implementation |
| CGinS-gui/benchmarks | Profiling and benchmark workspace |
| CGinS-gui/kernels | CUDA kernel workspace |
| CGinS-gui/cgins_runtime | Runtime integration |
| CGinS-gui/main.jac | Jac application entry point |
| CGinS-gui/cgins-frontend | Frontend implementation |
git clone https://github.com/Dhravidk/openTorch.git
cd openTorch/CGinS-guiStart with the implementation README and Jac guide for environment and provider setup. GPU execution requires a compatible NVIDIA/CUDA environment; generation uses a configured model provider.
This is an exploratory research implementation. Consult the included paper for its experimental setting and reported results. Performance depends on the operator, workload, hardware, and baseline; the presence of a generated kernel does not establish an application-level speedup.
The original implementation and its documentation remain under CGinS-gui/.