AI Platforms
Platforms for hosting, deploying, and running AI models
Hugging Face
Hugging Face
The largest open-source community for AI models and datasets, offering model hosting, inference APIs, and collaboration tools.
Replicate
Replicate
A cloud platform for running AI models, letting you run open-source models without managing infrastructure.
Ollama
Ollama
The standard tool for running open-source LLMs locally: one command downloads and serves models like Llama, Qwen and DeepSeek on your own machine, with a built-in API.
Together AI
Together AI
A cloud platform for training, fine-tuning and serving open-source models at scale, with fast inference APIs across hundreds of checkpoints.
Groq
Groq
An inference platform built on custom LPU hardware, famous for generating open-model responses at exceptional speeds through a simple API.
OpenRouter
OpenRouter
A unified API that routes requests to hundreds of models from every major lab, with price and latency comparisons and automatic fallbacks.
DeepInfra
DeepInfra
A pay-per-token inference cloud serving hundreds of open models — text, image, speech and embedding — through a simple standardized API.
Fireworks AI
Fireworks
A fast inference cloud for open models with production-grade reliability, fine-tuning and compositional multi-model deployments.
Baseten
Baseten
An inference platform for deploying custom and open models on dedicated GPUs, with autoscaling, low-latency serving and model privacy.
Modal
Modal Labs
Serverless cloud for AI and machine learning: run inference jobs, fine-tuning and batch workloads on autoscaling GPUs with code-defined infrastructure.
RunPod
RunPod
A GPU cloud for AI workloads: serverless inference, on-demand and spot GPUs, and pod environments for training and hosting open models.
Vast.ai
Vast.ai
A marketplace for renting community-run GPUs at the lowest prices, supporting training, inference and rendering workloads with flexible bidding.
Novita AI
Novita
An affordable inference cloud for hundreds of open models — LLMs, image and audio generation — with per-token and per-image pricing.
SiliconFlow
SiliconFlow
A China-rooted inference cloud serving leading open models — GLM, Qwen, DeepSeek and more — through standardized, low-cost APIs.
Perplexity Sonar API
Perplexity
Perplexity's search-grounded inference API: LLM responses with real-time web citations, built for applications that need factual, sourced answers.
Azure OpenAI
Microsoft
Microsoft Azure's enterprise service for OpenAI models: GPT-4o, o-series and embedding models with enterprise security, compliance and regional availability.
Cohere
Cohere
An enterprise LLM platform built for security-conscious organizations: the Command model family, retrieval-grade Embed and Rerank models, and private deployment on your own cloud or on-premises.
Writer
Writer
A full-stack enterprise generative AI platform: Palmyra LLMs, a no-code agent builder, Knowledge Graph grounding and governance controls for rolling out AI across regulated organizations.
vLLM
vLLM Project
The open-source inference engine behind much of modern LLM serving: PagedAttention delivers state-of-the-art throughput with an OpenAI-compatible API and broad model support on your own GPUs.
ModelScope
Alibaba (Dammo Academy)
Alibaba's open-source model community and MaaS platform: hundreds of thousands of models and datasets, free online inference notebooks, and the home hosting of the Qwen family.
Dify
Dify (LangGenius)
The open-source LLM app platform: build RAG assistants and agent workflows visually, back them with your choice of 100+ models, and run them on your own infrastructure or Dify Cloud.
Exa
Exa AI
A search engine built for AI: an API that returns clean, semantically relevant web content — page contents, not ad-laced result lists — designed to be read by models rather than humans.
Tavily
Tavily
A search API purpose-built for LLMs and RAG: one call returns ranked, content-ready web results optimized for retrieval — the plug-and-play web tool of countless agent frameworks.
Serper
Serper Dev
A fast, low-cost Google Search results API built for AI: structured SERP data (organic, news, images, maps) at a fraction of legacy SERP-API pricing — the workhorse behind many research agents.
fal.ai
fal
Generative-media cloud infrastructure: host and serve image, video and audio models through one API, famous for sub-second FLUX generation and day-one hosting of new open models.
Runware
Runware
Ultra-low-cost image generation infrastructure: a sub-second API claiming the industry's lowest per-image price by owning its GPU pipeline end to end — built for apps generating at massive volume.
Vapi
Vapi
The developer platform for voice agents: assemble sub-second phone-call AI from best-in-class STT, LLMs and TTS — with phone numbers, transfers, and integrations — without building the telephony plumbing yourself.
Brave Search API
Brave Software
An independent search index exposed as an API: privacy-first Google-alternative results with web, news, image and local endpoints — the independent-data option in the AI search stack.
Open WebUI
Open WebUI
The self-hosted ChatGPT-style interface for local and private models: a polished chat UI over Ollama, OpenAI-compatible APIs or any backend — with RAG, multi-user management and tool calling built in.
LibreChat
LibreChat
An open-source, multi-provider chat interface: one self-hosted UI for OpenAI, Anthropic, Google, local models and more — with agents, RAG, artifacts and multi-user management, no per-vendor lock-in.
Apple Intelligence
Apple
Apple's on-device and Private-Cloud AI layer: writing tools, notification summaries, image generation and a redesigned Siri — engineered so personal data never has to leave your device.
Pinecone
Pinecone
The managed vector database that defined the category: millisecond similarity search over billions of embeddings, serverless scaling, and the retrieval backbone of countless production RAG systems.