LM Studio
Discover, download, and run local LLMs on your Mac, Windows, or Linux machine completely offline.
About LM Studio
LM Studio is the premier desktop application for discovering, downloading, and running large language models locally on your personal computer. With a polished visual interface, LM Studio makes running models like Llama 3.3, DeepSeek-R1, Mistral, and Qwen as simple as a single click, keeping all personal conversations, code, and documents 100% private and offline on your hardware. Whether running on Apple Silicon with unified memory or dedicated Nvidia/AMD GPUs, LM Studio automatically configures optimal hardware acceleration (Metal, CUDA, ROCm, Vulkan) for smooth, high-speed local inference.
Beyond conversational chat, LM Studio includes a built-in local developer server that exposes OpenAI-compatible `/v1/models` and `/v1/chat/completions` endpoints on localhost. This enables developers to connect IDE extensions (Cursor, VS Code, Windsurf) and local agent pipelines directly to offline models. It features comprehensive support for GGUF model formats from Hugging Face, multi-model side-by-side benchmarking, GPU layer offloading controls, and system prompt presets.
Step 1: Download and install LM Studio for macOS, Windows, or Linux from lmstudio.ai.
Step 2: Search for trending models (e.g. Llama-3.3-8B, DeepSeek-R1-Distill) and click Download.
Step 3: Open the Chat tab, load the model into memory, and start private offline conversations.
Step 4: Optionally switch to the Local Server tab to expose a local OpenAI API endpoint on port 1234.
Capabilities & Features
Common Use Cases
Running 100% private, offline LLMs for confidential code and document analysis
Local OpenAI-compatible API server for Cursor, Windsurf, and agent development
Side-by-side benchmarking of quantized GGUF models on Apple Silicon and Nvidia GPUs
Exploring and testing newly released open-weights models from Hugging Face
Frequently Asked Questions
Does LM Studio send my chat data to the cloud?
No. LM Studio runs models entirely on your local CPU and GPU. No chat history, prompts, or model weights are ever sent to external cloud servers.
Can I connect Cursor or VS Code to LM Studio?
Yes. Start the local server in LM Studio (defaulting to http://localhost:1234/v1) and point your IDE API endpoint to it.
What hardware do I need for LM Studio?
LM Studio runs well on Apple Silicon Macs (M1/M2/M3/M4 with 16GB+ unified memory) or Windows/Linux PCs with Nvidia GeForce GPUs (8GB+ VRAM).
Free Plan
100% Free for personal use and developer experimentation
Paid Plan
LM Studio for Business licensing for enterprise teams and commercial redistribution
Direct link · Verified & reader-supported
Pros & Cons
Zero-configuration setup on Mac (Metal) and Windows/Linux (CUDA/Vulkan)
Integrated Hugging Face search to download GGUF models directly within the app
Local OpenAI-compatible API server powers Cursor and local AI workflows
100% private and offline—zero data leaves your local machine
Inference speed is bounded by your local hardware VRAM and RAM bandwidth
Core desktop application is proprietary (freeware for personal use)
Alternatives
View allOllama
Run powerful AI models locally on your machine.
Ollama allows users to run large language models directly on their local hardware, providing privacy and speed without relying on cloud services. It supports a variety of models and is optimized for both CPU and GPU usage. The tool is ideal for developers and researchers who need offline access to AI capabilities.
Hyperbolic
Decentralized GPU cloud & open-source AI inference engine
Hyperbolic is a decentralized AI computing network and inference platform that provides high-performance, cost-effective GPU compute and open-access LLM APIs. By aggregating global GPU infrastructure with cryptographic verification of compute, Hyperbolic delivers up to 75% cost savings on frontier open-source model inference compared to traditional cloud hyperscalers. Hyperbolic serves the fastest and most affordable APIs for open-source models including DeepSeek R1/V3, Llama 3.3 70B, Qwen 2.5, and SDXL with full OpenAI-compatible API endpoints.
RunPod
Globally distributed GPU cloud and serverless platform for AI inference and training
RunPod is a leading globally distributed GPU cloud and serverless computing platform engineered specifically for artificial intelligence workloads. It provides developers, AI researchers, and enterprises with on-demand access to top-tier NVIDIA GPUs (including H100, A100, L40S, and RTX 4090) at up to 80% lower cost than traditional legacy hyperscalers.
vLLM
High-throughput and memory-efficient LLM serving engine powered by PagedAttention.
vLLM is the industry-standard open-source LLM serving and inference engine designed for ultra-high throughput and minimal memory waste. Developed by UC Berkeley researchers, vLLM introduced PagedAttention—a revolutionary memory management algorithm that manages attention key-value (KV) cache like virtual memory in operating systems, virtually eliminating memory fragmentation. Capable of delivering 2x to 4x higher throughput than Hugging Face TGI and standard PyTorch runtimes, vLLM powers production AI inference infrastructure across enterprise cloud clusters and high-volume API providers worldwide.
Cursor Directory
Curated library of system prompts, framework templates, and .cursorrules for AI engineers
Cursor Directory is the leading open repository of specialized system prompts, framework rules, and `.cursorrules` configurations for Cursor IDE and modern AI coding assistants. It helps developers immediately optimize AI responses for specific technology stacks, from Next.js and Tailwind to Rust, LangChain, and Python backend microservices. Engineered by developers for developers, the platform categorizes community-tested rules by framework, language, and architectural design pattern. Instead of writing verbose context prompts on every task, developers can drop these `.cursorrules` directly into their repository root to enforce clean code conventions, typed schemas, and idiomatic practices effortlessly.
Raycast
Extensible AI-powered launcher for power users
Raycast replaces standard OS launchers with a command-driven interface for apps, scripts, and AI. It provides quick access to AI chat, translations, and custom workflows.
Compare LM Studio with Alternatives
Side-by-side feature, pricing, and pros & cons breakdowns
Featured in In-Depth Guides & Comparisons
Read our hands-on technical evaluations and workflow guides mentioning LM Studio



