Skip to content

fix(cuda-macros): correct gpu_printf default float and unsigned formatting - #1353

Open
YeonwooSung wants to merge 3 commits into
NVIDIA:mainfrom
YeonwooSung:fix/gpu-printf-default-format
Open

YeonwooSung wants to merge 3 commits into
NVIDIA:mainfrom
YeonwooSung:fix/gpu-printf-default-format

Conversation

@YeonwooSung

Copy link
Copy Markdown

What was wrong

gpu_printf!("{:.2}", 3.14159) is documented to print 3.14, but a placeholder with precision and no type character left format_type == None. That arm packed the argument as i64 and emitted %.2lld. 3.14159 as i64 is 3, and C precision on an integer conversion is a minimum digit count, so the result was 03, not 3.14.

The same default path cast every untyped or unsigned value with as i64 and printed it with %lld. A u64 above i64::MAX therefore wrapped and printed as a negative number.

This fixes an unfiled bug (no issue number).

Test that failed first

tests::printf::precision_without_type_on_float_uses_float_conversion failed before any production change. Expanding {:.2} on 3.14159f32 produced (3.14159f32) as i64 and b"%.2lld\0".

After that fix was in, tests::printf::default_format_on_u64_uses_unsigned_conversion failed on {} of 18446744073709551615u64, which expanded to (18446744073709551615u64) as i64 and b"%lld\0".

Fix

  • {:.N} with no type character uses %f and as f64. That includes arguments whose type is not visible in the tokens, so a float is not truncated to an integer.
  • A visible unsigned argument (42u64, value as u64) is packed with as u64 and printed with %llu.
  • A bare binding, whose type the macro cannot see, is packed through GpuPrintfArg. The format character comes from that trait (%llu for u64, %f for floats, %d / %lld for signed integers) instead of always guessing i64.
  • Explicit specifiers keep their previous conversions: {:x}, {:e}, {:f}, {:08}, and width on an unsuffixed integer.

How I verified

  • Red: the two tests above failed with as i64 / %.2lld and as i64 / %lld respectively.
  • Green: cargo test -p cuda-macros --lib (114 passed), including both new tests.
  • cargo test -p cuda-device --lib default_format_typechecks_for_visible_and_untyped_args typechecks the expanded gpu_printf! for {:.2} on f32, {} on u64::MAX, {:08}, {:x}, and {:e}.

cargo test -p cuda-macros (integration tests included) also builds cuda-bindings, which requires a CUDA 13 toolkit. No toolkit is installed in this environment (CUDA_HOME / CUDA_TOOLKIT_PATH unset, no /usr/local/cuda), so those dev-dependency tests were not run. The lib tests do not need the toolkit.

`{:.N}` has no type character, so the default arm packed the argument
as i64 and emitted %.Nlld. C precision on an integer conversion is a
minimum digit count, which turned 3.14159 into 03 instead of 3.14.

Signed-off-by: YeonwooSung <neos960518@gmail.com>
Default `{}` packed every argument as i64 and emitted %lld, so a u64
above i64::MAX printed as a negative number. Visible unsigned arguments
now cast to u64 and use %llu. Bindings whose type the macro cannot see
go through GpuPrintfArg, which also keeps floats off the integer path.

Signed-off-by: YeonwooSung <neos960518@gmail.com>
…printf-default-format

Signed-off-by: YeonwooSung <neos960518@gmail.com>

# Conflicts:
#	cuda-oxide/crates/cuda-macros/src/tests/printf.rs
@copy-pr-bot

copy-pr-bot Bot commented Oct 5, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant