llama.cpp local GGUF inference + HF Hub model discovery.
MLOps / AI 工程
218 skills
Track how 350+ AI models mention your brand — manage projects, run scans, analyze results
NeuralBrain MCP Server - RAG, Vector Memory, LLM Routing, Agent Identity, x402 Payments
MCP server for AiDataTaskRunner Panel (MT5) - Control data generation and ML model training via MCP
Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-…
Use this to test an LLM change (new prompt, new model, new retrieval) on real traffic before rolling it out to everyone. Trigger on "A/B test my prompt", "roll out a new model safely", "compare two prompts in production", "canary this change", "does this actually improve things for real users". Mea…
Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.
Structured domain knowledge for AI agents. 42x more accurate than RAG, 11x fewer tokens.
In-memory vector search with TF-IDF and cosine similarity. x402 micropayment.
Event-sourced world model for multi-LLM agents: propose, validate, and read a shared state.
Self-hostable agentic-AI LMS: catalog, RAG tutor, FSRS reviews, AI authoring, ingest.
Self-hostable agentic-AI LMS: catalog, RAG tutor, FSRS reviews, AI authoring, ingest.
130+ QA & dev tools for AI agents: prompt injection, RAG testing, VLM eval, guardrails. Free.
Autonomous AI agent infrastructure: S3 storage, vector DB, MCP hosting. Free. EU-hosted.
Dubstrata: Causal Financial & Narrative CDN providing Graph RAG & Deep Research Reports
Live LLM API pricing: token prices, comparisons, cheapest-model lookups. No key required.
Real-time NSE/BSE stock sentiment, news NLP, technical analysis & RAG search via MCP.
Persistent memory and vector search for AI agents. Hosted, OAuth-protected via GitHub/Google.
Protein analysis: ESM-2/ESMC embeddings, mutation scoring, landscape scans, ESMFold structure.
Parse PDF/Word/PPT/HTML to Markdown; tables as JSON, image extraction, RAG chunking, page ranges.
Zero model drift between AI agents. Session continuity for Claude, GPT, Gemini, and any LLM.
Zero model drift between AI agents. Session continuity for Claude, GPT, Gemini, and any LLM.
Live LLM API price + status radar across 11 providers, with public per-model price HISTORY.
Local semantic search — embedding-powered grep for files, zero external services.
Provision private AI model endpoints on dedicated GPUs (Llama, Qwen, Mistral). Pay per minute.
Provision private AI model endpoints on dedicated GPUs (Llama, Qwen, Mistral). Pay per minute.
LLM cost optimization. Audit your calls, prove savings, route to the cheapest quality model. BYOK.
Transform audio recordings into professional Markdown documentation with intelligent summaries using LLM integration
One key, 100+ models — chat with any LLM and generate video, images, speech. Free trial at 370.ai.
Jailbreak API-served LLMs using G0DM0D3 techniques — Parseltongue input obfuscation (33 techniques), GODMODE CLASSIC system prompt templates, ULTRAPLINIAN multi-model racing, encoding escalation, and Kheish-native prefill/system prompt integration. Use when a user wants to bypass safety filters, ja…