Skip to content

Grid computing in HEP: basic concepts - #200

Open
michmx wants to merge 2 commits into
mainfrom
193-grid-basic-concepts
Open

michmx wants to merge 2 commits into
mainfrom
193-grid-basic-concepts

Conversation

@michmx

@michmx michmx commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator

Adding a chapter on basic concepts of distributed computing in HEP, intended for beginners.

Resolves #193 .

@michmx
michmx requested a review from kjvbrt September 4, 2026 03:41
@michmx michmx changed the title 193 grid basic concepts Grid computing in HEP: basic concepts Sep 4, 2026
@davidlange6

Copy link
Copy Markdown

I had two general comments:

a) is there perhaps someplace upstream (hsf?) that could host the non-fcc specific pieces of this?

b) perhaps docs on dirac, rucio, grid certs (perhaps better tokens instead) should wait until they are relevant for most (or at least many) users? [eg, rucio is much more advanced in fcc than this doc suggests]

@michmx

michmx commented Sep 4, 2026

Copy link
Copy Markdown
Collaborator Author

@davidlange6 Thanks for the review!

a) Right now, no place comes to my mind. Some time ago we considered the posibility of writting HSF modules related to grid computing, but we gave up since it not trivial to make generic enough material for examples and exercises (PanDA vs DIRAC vs glideinWMS, flat vs hierarchical namespaces, VO policies, etc). And we assumed each experiment has its own tailored version.

b) Yes you are right, I was lazy not digging deeper on the status of Rucio adoption. I can update that part. And about waiting that it becomes relevant for most users, the logic is that we already have quite some material about grid computing and we want to put some context for those not familiar with grid. I see no harm on adding it right now.

For example, CERN runs [HTCondor](https://batchdocs.web.cern.ch/) on its farm, reachable from `lxplus`. Many universities run HTCondor or Slurm. In batch system
you already have an account, the shared filesystem (AFS, EOS) is mounted on the worker nodes, and you can `ssh` in to look at things.

The grid glues many batch systems together behind a single interface. In exchange for scale, there is no shared filesystem across sites,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is CVMFS on all the sites as the one shared filesystem.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good point. I meant filesystem for users to write, but certainly better to be more accurate about it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Explanation of basic concepts in distributed computing

3 participants