Skip to content
View tommasocerruti's full-sized avatar
🔒
locked-in
🔒
locked-in

Highlights

  • Pro

Organizations

@evaleval

Block or report tommasocerruti

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
tommasocerruti/README.md

Pinned Loading

  1. harbor-framework/terminal-bench harbor-framework/terminal-bench Public

    Measuring and evolving with the frontier of agent work

    Python 620 444

  2. evaleval/every_eval_ever evaleval/every_eval_ever Public

    Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to loca…

    Python 109 48

  3. linear-attention-architectures linear-attention-architectures Public

    Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing (Technical Report)

    Python 35

  4. detllm detllm Public

    Deterministic-mode checks for LLM inference: measure run/batch variance, generate repro packs, and explain why outputs differ.

    Python 20 1

  5. algolab-2024 algolab-2024 Public

    Algorithms Lab (Algolab) HS 2024 @ ETH Zurich, solutions and insights

    C++ 17 3

  6. eal-bench eal-bench Public

    EAL-Bench: Benchmarking Endogenous Authorization Laundering (EAL) in Agent Memory

    Python 9 1