Skip to content

geom_spidergram (new) - #53

Draft
archaeothommy wants to merge 19 commits into
mainfrom
geom-spidergram
Draft

archaeothommy wants to merge 19 commits into
mainfrom
geom-spidergram

Conversation

@archaeothommy

@archaeothommy archaeothommy commented Jul 3, 2026 •

Copy link
Copy Markdown
Owner

New approach to geom_spidergram(), allowing to supply a character vector of elements as aesthetic.

Previous code for data treatment is copied at the end of the file.

@archaeothommy

Copy link
Copy Markdown
Owner Author

I was not able to fix the "discrete values to continuous scale" problem other than the solution you suggested. I have no idea why this occurs. We create the x aesthetic in setup data and it is a discrete aesthetic from the onset... I don't understand why ggplot2 thinks this is a continuous variable and never had this problem in the other geoms.

The only difference I see is how the graphical elements are created; the coordinate transformation might be the issue. There are two differences in your approach compared to others:

  • e.g. this case study relies on the functions of existing geoms.
  • My other geoms plot from data object, rather than coord, without transformation.

And for whatever reason, the geom doesn't find the elements_data and other pre-compiled objects...

Comment thread R/ASTR_normalise_wrapper.R Outdated
normalise_geochem(df, reference = reference, ...)
},
hundred = {
numeric_cols <- names(df)[sapply(df, is.numeric)]

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This test should be moved to the normalisation function so that this function does not fail if it is called outside of the wrapper function.

Comment thread R/ASTR_normalise_wrapper.R Outdated
df[numeric_cols] <- normalise_rows(df[numeric_cols])
df
},
element = {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please create a separate function for this normalisation. The wrapper is a convenience function and should not do anything else than dispatching it to functions that do the actual work.

Comment thread R/ASTR_normalise_wrapper.R Outdated
df[numeric_cols] <- lapply(df[numeric_cols], function(x) x / divisor)
df
},
sample = {

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please create a separate function for this normalisation which is them called here. The wrapper is a convenience function and should not do anything else than dispatching it to functions that do the actual work.

Comment thread R/ASTR_normalise_wrapper.R Outdated
Comment on lines +109 to +118
unknown = {
stop(
"Unknown reference '", reference, "'. ",
"`reference` must be one of: ",
"a geochemical reference composition (see `references_geochem`), ",
"a column name in `df`, ",
"an ID value in `df`, ",
"or '100%'."
)
}

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
unknown = {
stop(
"Unknown reference '", reference, "'. ",
"`reference` must be one of: ",
"a geochemical reference composition (see `references_geochem`), ",
"a column name in `df`, ",
"an ID value in `df`, ",
"or '100%'."
)
}
stop(
"Unknown reference '", reference, "'. ",
"`reference` must be one of: ",
"a geochemical reference composition (see `references_geochem`), ",
"a column name in `df`, ",
"an ID value in `df`, ",
"or '100%'."
)

The default function of switch statement is unnamed. It is automatically called in the input value does not match any other options.

@archaeothommy

Copy link
Copy Markdown
Owner Author
  • normalise_element: should state that the ratio between the element and the reference element is calculated
  • normalise_sample: check if provided is ID is unique as part of tests, not further down.

@karan3242

Copy link
Copy Markdown
Collaborator

@Abagna123 Could you let me know where the issue of log scale mismatch is happening?

Comment thread data/elements_data.rda

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Abagna123 I I am having a build error for this.
'elements_data' is not an exported object from 'namespace:ASTR'
Why were these and other data sets modified?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@karan3242 Nothing is missing or deleted. The data objects are internal by design. The build error comes from a single incorrect line in GeomSpider that uses ASTR::elements_data, which is not valid for internal package data.

optional_aes = c(ASTR::elements_data, ASTR::isotopes_data, ASTR::oxides_data),

  • elements_data, isotopes_data, oxides_data, conversion_oxides are created in the data prep script (in data-raw/).
  • They are saved with internal = FALSE, so they live in data/*.rda.
  • They are documented in R/ASTR_data.R, but they are not exported (NAMESPACE has no export() lines for them).

That part was done by Thomas so I am confused, if they were never supposed to be exported, then ASTR:: should never have been used on them.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Afaik exporting data does not work with the same @export -> NAMESPACE workflow as for functions. Hadley says "Never @export a data set.". But this also fails when you remove the ASTR::, which should usually work in functions inside a package. I think there is something weird about ggproto and when exactly it is evaluated.

After looking into this for a while I wonder if there even is a way to use this data here, if we want it to be available externally. Maybe the latter is not necessary? Then we could try the internal data setup lined out here: https://r-pkgs.org/data.html#sec-data-sysdata. Or we just have the data in there twice, internally and externally. A bit annoying, but this is an edge case.

@Abagna123

Copy link
Copy Markdown
Collaborator

@karan3242 library(ggplot2)

test_df <- data.frame(
Sample = c("A", "B", "C"),
La = c(10, 5, 8),
Ce = c(20, 8, 15),
Pr = c(5, NA, 4),
Nd = c(15, 6, 12),
Sm = c(8, 3, NA),
Eu = c(NA, NA, 5),
Gd = c(12, 4, 9),
Tb = c(2, 1, NA),
Dy = c(8, 3, 6),
Ho = c(2, 1, 2),
Er = c(5, 2, 4),
Tm = c(1, NA, 1),
Yb = c(4, 2, 3),
Lu = c(1, 1, 1)
)

elements <- c("La", "Ce", "Pr", "Nd", "Sm", "Eu", "Gd", "Tb", "Dy", "Ho", "Er", "Tm", "Yb", "Lu")

1. Raw values, linear scale

ggplot(test_df) +
geom_spider(
aes(elements = elements, colour = Sample)
) +
scale_y_continuous() +
labs(title = "Raw values, linear scale", x = NULL, y = "Concentration (ppm)") +
theme_bw()

2. Chondrite normalised, log scale

ggplot(test_df) +
geom_spider(
aes(elements = elements, colour = Sample),
reference = "chondrite"
) +
scale_y_log10() +
labs(title = "Chondrite normalised, log scale",
x = NULL, y = "Sample / Chondrite") +
theme_bw()


Test 1 and Test 2 are identical except for the y scale. Test 1 Lines plot correctly, y-axis shows normal values. Test 2 Y-axis shows 1e+25, 1e+42.... and even without y axis values when log_scale is applied . So i suspected geom_spider() creates the y values inside setup_data(). But ggplot2 applies the y-scale transform earlier in the build pipeline. So the transform runs before y exists. The axis then gets log labels while the data stays raw (e.g 1e+25, 1e+42).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants