SKILL_MD
24

Tests an agent's capacity to handle complex constraint satisfaction by managing and resolving conflicting schedules derived purely from raw natural language descriptions. It simulates personal assistant scenarios where the model must maintain a high-density representation of events to detect confli…

Productivity / 办公效率Safe-ishScanned
SKILL_MD
24

This benchmark evaluates text anonymization methods by measuring both span-level masking accuracy and subject-level privacy leakage. It probes whether anonymized text successfully prevents adversarial LLMs from inferring personal identifiable information (PII) and sensitive attributes, while mainta…

MLOps / AI 工程Safe-ishScanned
SKILL_MD
24

Evaluates a general-purpose NeRF framework on downstream 3D tasks, including real-time novel view synthesis, parameter-efficient 3D scene understanding, and text-guided 3D editing. It probes the model's ability to generalize to unseen scenes and adapt to various geometric and appearance tasks with …

Creative / 创意娱乐Safe-ishScanned
SKILL_MD
binarystatscores
by qhjqhj00
24

Compute the BinaryStatScores metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute BinaryStatScores, or asks how to score with BinaryStatScores.

Safe-ishScanned
SKILL_MD
24

Evaluates real-time urban pathfinding algorithms under dynamic traffic and weather conditions. It measures how well traditional graph search methods and deep learning models predict optimal routes and minimize travel time in a simulated Berlin city environment. Use when the user wants to benchmark …

MLOps / AI 工程Safe-ishScanned
24

Evaluates a unified framework's ability to generate realistic and semantically aligned 3D human motions from text descriptions, recognize actions from skeleton data, and retrieve matching text-motion pairs. It probes the model's semantic fidelity, distributional realism, and cross-modal alignment c…

Coding / 软件开发Safe-ishScanned
SKILL_MD
24

This benchmark probes large language models' ability to comprehend implicit code review intent by decomposing the task into change type recognition, change localization, and solution identification. It uses multiple-choice questions to evaluate whether models can accurately interpret pre-change cod…

Coding / 软件开发Safe-ishScanned
SKILL_MD
chikha-po-eval
by qhjqhj00
24

Evaluates fundamental lexical comprehension and generation capabilities of multilingual LLMs across thousands of languages. It probes word-level translation, context-aware translation, translation-conditioned language modeling, and bag-of-words machine translation tasks to measure basic linguistic …

Writing / 写作Safe-ishScanned
SKILL_MD
atlas-chat-eval
by qhjqhj00
24

Evaluates large language models on Moroccan Arabic (Darija) across multiple-choice reasoning, instruction following, translation, summarization, and sentiment analysis. It probes dialect-specific linguistic features, script variability (Arabic vs. Arabizi), and real-world instruction-following capa…

Writing / 写作Safe-ishScanned