Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect confli…
qhjqhj00
9 skills Unclaimed
This benchmark evaluates text anonymization methods by measuring both span-level masking accuracy and subject-level privacy leakage. It probes whether anonymized text successfully prevents adversarial LLMs from inferring personal identifiable information (PII) and sensitive attributes, while mainta…
Evaluates a general-purpose NeRF framework on downstream 3D tasks, including real-time novel view synthesis, parameter-efficient 3D scene understanding, and text-guided 3D editing. It probes the model's ability to generalize to unseen scenes and adapt to various geometric and appearance tasks with …
Compute the BinaryStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryStatScores, or asks how to score with BinaryStatScores.
Evaluates real-time urban pathfinding algorithms under dynamic traffic and weather conditions. It measures how well traditional graph search methods and deep learning models predict optimal routes and minimize travel time in a simulated Berlin city environment. Use when the user wants to benchmark …
Evaluates a unified framework's ability to generate realistic and semantically aligned 3D human motions from text descriptions, recognize actions from skeleton data, and retrieve matching text-motion pairs. It probes the model's semantic fidelity, distributional realism, and cross-modal alignment c…
This benchmark probes large language models' ability to comprehend implicit code review intent by decomposing the task into change type recognition, change localization, and solution identification. It uses multiple-choice questions to evaluate whether models can accurately interpret pre-change cod…
Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic …
Evaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capa…