{"id":42964,"date":"2026-02-09T01:00:00","date_gmt":"2026-02-09T09:00:00","guid":{"rendered":"https:\/\/dhblog.dream.press\/blog\/?p=42964"},"modified":"2026-08-03T21:48:15","modified_gmt":"2026-08-04T04:48:15","slug":"open-source-ai","status":"publish","type":"post","link":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/","title":{"rendered":"The 10 Best Self-Hosted AI Models You Can Run at Home"},"content":{"rendered":"<div class=\"tldr-block\" style=\"display: none;\">\n\t<div class=\"svg\">\n\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" viewBox=\"0 0 119.25 37.8\">\n\t<g>\n\t\t<g>\n\t\t\t<path fill=\"#ffffff\" d=\"M23.4,6.93h-8.1v24.57h-7.2V6.93H0V0h23.4v6.93Z\" \/>\n\t\t\t<path fill=\"#ffffff\" d=\"M45,24.57v6.93h-18.45V0h7.2v24.57h11.25Z\" \/>\n\t\t\t<path fill=\"#ffffff\"\n\t\t\t\td=\"M90.9,15.75c0,8.91-6.61,15.75-15.3,15.75h-12.6V0h12.6c8.68,0,15.3,6.84,15.3,15.75ZM83.97,15.75c0-5.4-3.42-8.82-8.37-8.82h-5.4v17.64h5.4c4.95,0,8.37-3.42,8.37-8.82Z\" \/>\n\t\t\t<path fill=\"#ffffff\"\n\t\t\t\td=\"M105.57,21.15h-3.42v10.35h-7.2V0h12.6c5.98,0,10.8,4.81,10.8,10.8,0,3.87-2.34,7.38-5.81,9.13l6.71,11.56h-7.74l-5.94-10.35ZM102.15,14.85h5.4c1.98,0,3.6-1.75,3.6-4.05s-1.62-4.05-3.6-4.05h-5.4v8.1Z\" \/>\n\t\t<\/g>\n\t\t<path\n\t\t\tfill=\"#0173ec\"\n\t\t\td=\"M53.97,37.8h-5.4l1.8-13.27h7.2l-3.6,13.27ZM49.02,12.55c0-2.34,1.93-4.27,4.27-4.27s4.27,1.94,4.27,4.27-1.93,4.27-4.27,4.27-4.27-1.94-4.27-4.27Z\"\n\t\t \/>\n\t<\/g>\n<\/svg>\n\t<\/div>\n\t<div class=\"tldr-wrap\">\n\t\t\n\n<p class=\"wp-block-paragraph\">Most of the \u201copen source\u201d AI models are actually \u201copen-weight,\u201d which enables local, API-free use. If you want to run more powerful models, you need to use quantization, which can reduce model size by about 75%.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The hardware you need for local AI at a minimum:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>8GB VRAM: Entry-level 3B-7B models (e.g., Ministral).<\/li>\n\n\n\n<li>12GB VRAM: Daily use 8B models (e.g., Qwen3).<\/li>\n\n\n\n<li>16GB VRAM: Complex 14B-20B models (e.g., Phi-4, gpt-oss).<\/li>\n\n\n\n<li>24GB+ VRAM: Power users.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Use Ollama (open source, easy setup) or LM Studio (free but closed-source GUI) for deployment. Local AI is best suited to single users. Team access and guaranteed uptime require dedicated server infrastructure.<\/p>\n\n\n\t<\/div>\n<\/div>\n\n\n<p class=\"wp-block-paragraph\"><strong>The short answer: Ministral 3 8B is the best self-hosted AI model for 12GB GPUs, OpenAI\u2019s gpt-oss-20b for 16GB, and Qwen3 VL 32B for 24GB+ (all Apache 2.0 licensed and covered below).<\/strong> Licensing matters more than you\u2019d think, because half the \u201copen-source\u201d models people recommend on Reddit would make Richard Stallman\u2019s eye twitch. Llama uses a Community license with strict usage restrictions, and Gemma models through Gemma 3 came with Terms of Use you should <em>absolutely <\/em>read before shipping anything with them (Gemma 4 moved to Apache 2.0 in April 2026).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The term itself has become meaningless due to overuse, so before we recommend any software, let\u2019s first clarify the definition.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What you actually need are open-weight models. Weights are the downloadable \u201cbrains\u201d of the AI. While the training data and methods might remain a trade secret, you get the part that matters: a model that runs entirely on hardware you control.<\/p>\n\n\n\n<div class=\"liftoff-cta-card\">\n\t<div class=\"line\">\n\t\t<svg width=\"834\" height=\"469\" viewBox=\"0 0 834 469\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n<path opacity=\"0.7\" d=\"M865.739 -59.8017C865.739 -59.8017 832.792 123.045 744.31 182.376C655.829 241.707 562.417 143.097 474.164 202.767C386.505 262.036 442.275 384.659 354.504 443.76C266.434 503.061 98.0198 364.278 4.7754 318.308\" stroke=\"url(#paint0_linear_8_19)\" stroke-opacity=\"0.25\" stroke-width=\"19.8\"\/>\n<defs>\n<linearGradient id=\"paint0_linear_8_19\" x1=\"918.374\" y1=\"-112.088\" x2=\"147.486\" y2=\"548.265\" gradientUnits=\"userSpaceOnUse\">\n<stop offset=\"0.0576923\"\/>\n<stop offset=\"0.350962\" stop-color=\"#0073EC\"\/>\n<stop offset=\"0.825067\" stop-color=\"#C265FE\"\/>\n<stop offset=\"1\"\/>\n<\/linearGradient>\n<\/defs>\n<\/svg>\n\n\t<\/div>\n\t<div class=\"liftoff-cta-card__content\">\n\t\t<div class=\"headline_1\">\n\t\t\t\n<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"32\" height=\"32\" viewBox=\"0 0 32 32\" fill=\"none\">\n<path d=\"M24.0006 16.0019C19.5835 16.0019 16.0015 19.5839 16.0015 24.001V32.0001H32.003V15.9985H24.0039L24.0006 16.0019Z\" fill=\"url(#paint0_linear_3747_604)\"\/>\n<path d=\"M16.0015 7.99911V0H0V16.0016H7.99906C12.4162 16.0016 15.9981 12.4196 15.9981 8.00247L16.0015 7.99911Z\" fill=\"url(#paint1_linear_3747_604)\"\/>\n<path d=\"M7.99902 16.002C12.4168 16.002 15.998 19.5832 15.998 24.001C15.9979 28.4186 12.4167 32 7.99902 32C3.58137 32 0.000149208 28.4186 0 24.001C0 19.5832 3.58128 16.002 7.99902 16.002ZM24.001 0C28.4185 0.000241972 32 3.58143 32 7.99902C32 12.4167 28.4185 15.9978 24.001 15.998C19.5832 15.998 16.002 12.4168 16.002 7.99902C16.002 3.58128 19.5832 0 24.001 0Z\" fill=\"url(#paint2_linear_3747_604)\"\/>\n<rect x=\"8\" y=\"8\" width=\"16\" height=\"16\" fill=\"#FFFFFF\"\/>\n<path d=\"M16.0015 7.99902H24.0006V15.9981C19.5835 15.9981 16.0015 12.4128 16.0015 7.99902Z\" fill=\"#18181B\"\/>\n<path d=\"M7.99908 16.0015L7.99908 8.00235H15.9981C15.9981 12.4195 12.4128 16.0015 7.99908 16.0015Z\" fill=\"#18181B\"\/>\n<path d=\"M16.0015 24.0005H8.00246V16.0014C12.4196 16.0014 16.0015 19.5867 16.0015 24.0005Z\" fill=\"#18181B\"\/>\n<path d=\"M24.0007 16.0015V24.0006H16.0016C16.0016 19.5835 19.5869 16.0015 24.0007 16.0015Z\" fill=\"#18181B\"\/>\n<defs>\n<linearGradient id=\"paint0_linear_3747_604\" x1=\"16.0001\" y1=\"16.0002\" x2=\"32.0001\" y2=\"32.0002\" gradientUnits=\"userSpaceOnUse\">\n<stop stop-color=\"#A1A1AA\"\/>\n<stop offset=\"1\" stop-color=\"#C7C7CD\"\/>\n<\/linearGradient>\n<linearGradient id=\"paint1_linear_3747_604\" x1=\"0\" y1=\"0\" x2=\"16\" y2=\"16\" gradientUnits=\"userSpaceOnUse\">\n<stop offset=\"0.251049\" stop-color=\"#C7C7CD\"\/>\n<stop offset=\"1\" stop-color=\"#A1A1AA\"\/>\n<\/linearGradient>\n<linearGradient id=\"paint2_linear_3747_604\" x1=\"-11.3782\" y1=\"44.9411\" x2=\"-8.40086\" y2=\"-18.7449\" gradientUnits=\"userSpaceOnUse\">\n<stop stop-color=\"#BE59FF\"\/>\n<stop offset=\"0.19\" stop-color=\"#9D60FF\"\/>\n<stop offset=\"0.74\" stop-color=\"#4274FF\"\/>\n<stop offset=\"1\" stop-color=\"#1F7CFF\"\/>\n<\/linearGradient>\n<\/defs>\n<\/svg>\n\n\t\t\tMeet Remixer\n\t\t<\/div>\n\t\t<div class=\"headline_2\">You describe it. Remixer builds it.<\/div>\n\t\t<p>The AI website builder that turns conversation into designer-level sites. Free with hosting.<\/p>\n\t\t        <a\n            href=\"https:\/\/www.dreamhost.com\/remixer-website-builder\/\"\n                        class=\"btn btn--brand\"\n                                    target=\"_blank\"\n            rel=\"noopener noreferrer\"\n            >\n                            Start Free Trial                    <\/a>\n\n\t<\/div>\n\t<div class=\"tr-img-wrap-outer\"><img decoding=\"async\" data-skip-lazy class=\"\" src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/03\/remixer-screen.webp\" alt=\"DreamHost Remixer AI website builder\" \/><\/div>\n<\/div>\n\n\n<h2 id=\"h2_what-is-self-hosted-ai\" class=\"wp-block-heading\">What is self-hosted AI?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Self-hosted AI means running an AI model on infrastructure you administer \u2014 a desktop GPU, a home server, or a rented dedicated server \u2014 rather than calling a third-party model API. You download an open-weight model once and run it yourself, with no per-token fees, no rate limits, and prompts and outputs staying within the environment you control.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Here are our top picks at a glance \u2014 every model gets a full breakdown below:<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Model<\/strong><\/td><td><strong>VRAM tier<\/strong><\/td><td><strong>License<\/strong><\/td><td><strong>Best for<\/strong><\/td><\/tr><tr><td>Ministral 3 8B<\/td><td>12GB<\/td><td>Apache 2.0<\/td><td>Fast all-purpose chat + images<\/td><\/tr><tr><td>Qwen3 8B<\/td><td>12GB<\/td><td>Apache 2.0<\/td><td>Agentic and tool-calling work<\/td><\/tr><tr><td>Llama 3.1 8B Instruct<\/td><td>12GB<\/td><td>Llama Community<\/td><td>Tutorial and tooling compatibility<\/td><\/tr><tr><td>Qwen3-Coder 30B (MoE)<\/td><td>12\u201316GB<\/td><td>Apache 2.0<\/td><td>Coding<\/td><\/tr><tr><td>Ministral 3 14B<\/td><td>16GB<\/td><td>Apache 2.0<\/td><td>Reliable reasoning<\/td><\/tr><tr><td>Phi-4 14B<\/td><td>16GB<\/td><td>MIT<\/td><td>Worry-free commercial use<\/td><\/tr><tr><td>gpt-oss-20b<\/td><td>16GB<\/td><td>Apache 2.0<\/td><td>Reasoning speed<\/td><\/tr><tr><td>Gemma 4 12B<\/td><td>16GB<\/td><td>Apache 2.0<\/td><td>Multimodal (text, images, audio)<\/td><\/tr><tr><td>Qwen3 VL 32B<\/td><td>24GB+<\/td><td>Apache 2.0<\/td><td>Vision + language at scale<\/td><\/tr><tr><td>Gemma 4 26B<\/td><td>24GB+<\/td><td>Apache 2.0<\/td><td>Near-frontier generalist<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h2 id=\"h-what-s-the-difference-between-open-source-open-weights-and-terms-based-ai\" class=\"wp-block-heading\">What\u2019s the difference between open-source, open-weights, and terms-based AI?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u201cOpen\u201d is a spectrum in modern AI that requires careful navigation to avoid legal pitfalls.<\/strong><\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"1600\" height=\"813\" data-src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x.webp\" alt=\"Horizontal comparison chart of open-source, open-weights\" class=\"wp-image-79439 lazyload\" data-srcset=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x.webp 1600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-300x152.webp 300w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-1024x520.webp 1024w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-768x390.webp 768w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-1536x780.webp 1536w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-600x305.webp 600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-1200x610.webp 1200w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-730x371.webp 730w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-1460x742.webp 1460w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-784x398.webp 784w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-1568x797.webp 1568w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/01-Open-Source-vs.-Open-Weights-vs.-Terms-Based-AI_1x-877x446.webp 877w\" data-sizes=\"(max-width: 1600px) 100vw, 1600px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1600px; --smush-placeholder-aspect-ratio: 1600\/813;\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">These three categories tell you what you\u2019re actually downloading and what you\u2019re allowed to do with it.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>Category<\/strong><\/td><td><strong>Definition<\/strong><\/td><td><strong>Typical Licenses<\/strong><\/td><td><strong>Commercial Safety<\/strong><\/td><\/tr><tr><td>Open Source AI (Strict)<\/td><td>Meets the <a target=\"_blank\" href=\"https:\/\/opensource.org\/ai\">Open Source Initiative (OSI) AI definition<\/a>; you get the weights, the code, and detailed information about the training data, which is the \u201cpreferred form\u201d to modify the model.<\/td><td>OSI-Approved<\/td><td>Absolute; you have total freedom to use, study, modify, and share.<\/td><\/tr><tr><td>Open-Weights<\/td><td>You can download and run the \u201cbrain\u201d (weights) locally, but the training data and recipe often remain closed.<\/td><td>Apache 2.0, MIT<\/td><td>High; generally safe for commercial products, fine-tuning, and redistribution.<\/td><\/tr><tr><td>Source-Available\/Terms-Based<\/td><td>Weights are downloadable, but specific legal terms strictly dictate how, where, and by whom they can be used.<\/td><td>Llama Community, Gemma Terms (Gemma 3 and earlier)<\/td><td>Restricted; often includes usage thresholds (e.g., &gt;700M users) and acceptable use policies.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<h3 class=\"wp-block-heading\">Why does the definition of \u201copen\u201d matter?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Open-weights models entered a more mature phase somewhere around mid-2025. \u201cOpen\u201d increasingly means not just downloadable weights, but how much of the system you can <a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/blog\/local-ai-hosting\/\">inspect, reproduce, and govern<\/a>.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Open is a spectrum:<\/strong> In AI, \u201copen\u201d isn\u2019t a yes\/no label. Some projects open weights, others open training recipes, and others open evaluations. The more of the stack you can inspect and reproduce, the more open it really is.<\/li>\n\n\n\n<li><strong>The point of openness is <\/strong><a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/blog\/data-portability\/\"><strong>sovereignty<\/strong><\/a><strong>:<\/strong> The real value of open-weight models is their control. You can run them where your data lives, tune them to your workflows, and keep operating even when vendors change pricing or policies.<\/li>\n\n\n\n<li><strong>Open means auditable:<\/strong> Openness doesn\u2019t magically remove bias or hallucinations, but what it does give you is the ability to audit the model and apply your own guardrails.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">\ud83d\udca1<strong>Pro tip:<\/strong> If you\u2019re unsure what category the model you picked falls into, do a quick sanity check. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/docs\/hub\/en\/model-cards\">Find the model card on Hugging Face<\/a>, scroll to the license section, and read it. Apache 2.0 is usually the safest choice for commercial deployment.<\/p>\n\n\n\n<div class=\"article-newsletter article-newsletter--gradient\">\n\n\n<h2>Get Content Delivered Straight to Your Inbox<\/h2><p>Subscribe now to receive all the latest updates, delivered directly to your inbox.<\/p><form class=\"nwsl-form\" id=\"newsletter_block_\" novalidate><div class=\"messages\"><\/div><div class=\"form-group\"><label for=\"input_newsletter_block_\"><input type=\"email\"name=\"email\"id=\"input_newsletter_block_\"placeholder=\"Enter your email address\"novalidatedisabled=\"disabled\"\/><\/label><button type=\"submit\"class=\"btn btn--brand\"disabled=\"disabled\"><span>Sign Me Up!<\/span><svg width=\"21\" height=\"14\" viewBox=\"0 0 21 14\" fill=\"none\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\">\n<path d=\"M13.8523 0.42524L12.9323 1.34521C12.7095 1.56801 12.7132 1.9304 12.9404 2.14865L16.7241 5.7823H0.5625C0.251859 5.7823 0 6.03416 0 6.3448V7.6573C0 7.96794 0.251859 8.2198 0.5625 8.2198H16.7241L12.9405 11.8535C12.7132 12.0717 12.7095 12.4341 12.9323 12.6569L13.8523 13.5769C14.072 13.7965 14.4281 13.7965 14.6478 13.5769L20.8259 7.39879C21.0456 7.17913 21.0456 6.82298 20.8259 6.60327L14.6477 0.42524C14.4281 0.205584 14.0719 0.205584 13.8523 0.42524Z\" fill=\"white\"\/>\n<\/svg>\n<\/button><\/div><\/form><\/div>\n\n\n<h2 id=\"h2_how-does-gpu-memory-determine-which-models-you-can-run\" class=\"wp-block-heading\">How does GPU memory determine which models you can run?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Nobody chooses the \u201cbest\u201d model on the market. People choose the model that best fits their VRAM without crashing. The benchmarks are irrelevant if a model requires 48GB of memory and you are running an RTX 4060.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">To avoid wasting time on testing impossible recommendations, here are three distinct factors that consume your GPU memory during inference:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Model weights:<\/strong> This is your baseline cost. An 8-billion parameter model at full precision (FP16) needs roughly 16GB just to load \u2014 double the parameters, double the memory.<\/li>\n\n\n\n<li><strong>Key-value cache:<\/strong> This grows with every word you type. Every token processed allocates memory for \u201cattention,\u201d meaning a model that loads successfully might still crash halfway through a long document if you max out the context window.<\/li>\n\n\n\n<li><strong>Overhead:<\/strong> Frameworks and CUDA drivers permanently reserve another 0.5GB to 1GB. This is non-negotiable, and that memory is simply gone.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">However, if you want to run larger parameter models, look into quantization. <strong>Quantizing the weight precision from 16-bit to 4-bit can shrink a model\u2019s footprint by roughly 75% with barely any loss in quality.<\/strong><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The industry standard Q4_K_M (GGUF format) keeps chat and coding quality close to the full-precision original while reducing the memory requirements.<\/p>\n\n\n\n<h2 id=\"h2_what-can-you-expect-from-different-vram-configurations\" class=\"wp-block-heading\">What can you expect from different VRAM configurations?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Your VRAM tier dictates your experience, from fast, simple chatbots to near-frontier reasoning capabilities. This quick table is a realistic look at what you can run.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><tbody><tr><td><strong>GPU VRAM<\/strong><\/td><td><strong>Comfortable Model Size (Quantized)<\/strong><\/td><td><strong>What to Expect<\/strong><\/td><\/tr><tr><td>8GB<\/td><td>~3B to 7B parameters<\/td><td>Fast responses, basic coding assistance, and simple chat.<\/td><\/tr><tr><td>12GB<\/td><td>~7B to 10B parameters<\/td><td>The \u201cDaily Driver\u201d sweet spot; solid reasoning, good instruction following.<\/td><\/tr><tr><td>16GB<\/td><td>~14B to 20B parameters<\/td><td>A noticeable capability jump; better code generation and complex logic.<\/td><\/tr><tr><td>24GB+<\/td><td>~27B to 32B parameters<\/td><td>Near-frontier quality; slower generation, but great for RAG and long documents.<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\ud83e\udd13Nerd note:<\/strong> Context length can blow up memory faster than you expect. A model that runs fine with 4K context might choke at 32K. So, don\u2019t max out context unless you\u2019ve done the math.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>No NVIDIA card? The tiers still apply \u2014 just read them against your usable memory.<\/strong> On Apple Silicon, the GPU shares unified memory with everything else on the machine (check yours under Apple menu &gt; About This Mac), so budget conservatively: a 16GB Mac lives in the 8GB tier, and a 32GB Mac can work in the 16GB tier. Ollama runs natively on Apple Silicon, and LM Studio splits models between CPU and GPU automatically. AMD GPUs work too: Ollama ships an AMD path on Linux, and the Docker-based stacks covered below include an AMD profile.<\/p>\n\n\n\n<h2 id=\"h2_the-10-best-self-hosted-ai-models-you-can-run-at-home\" class=\"wp-block-heading\">The 10 best self-hosted AI models you can run locally<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">We\u2019re grouping these by VRAM tier because that is what actually matters. Benchmarks come and go, but your GPU\u2019s memory capacity is a physical constant.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you think in tasks rather than gigabytes, cross-reference this way:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Coding:<\/strong> Qwen3-Coder 30B (MoE) \u2014 the current dedicated coder (#4 below).<\/li>\n\n\n\n<li><strong>Private document and image analysis:<\/strong> Ministral 3 8B or 14B, Gemma 4 12B, or, at 24GB+, Qwen3 VL 32B; all are multimodal.<\/li>\n\n\n\n<li><strong>Smart-home and automation tool calling (think Home Assistant):<\/strong> a small, fast tool-caller like Qwen3 4B. At 4-bit quantization that\u2019s only about 2GB of weights (4B \u00d7 half a byte), so responses stay snappy on hardware that would choke on a reasoning model.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">Best self-hosted AI models for 12GB VRAM<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For the 12GB tier, you are looking for efficiency. You want models that punch above their weight class.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" width=\"1202\" height=\"1309\" data-src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image.png.webp\" alt=\"Grid of four AI model cards for 12GB VRAM \u2014 Ministral 3 8B, Qwen3 8B, Llama 3.1 8B Instruct, and Qwen3-Coder 30B (MoE)\" class=\"wp-image-86678 lazyload\" data-srcset=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image.png.webp 1202w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-275x300.png 275w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-940x1024.png 940w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-768x836.png 768w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-600x653.png.webp 600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-1200x1307.png.webp 1200w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-730x795.png.webp 730w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-784x854.png.webp 784w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-12gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-0-image-877x955.png.webp 877w\" data-sizes=\"(max-width: 1202px) 100vw, 1202px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1202px; --smush-placeholder-aspect-ratio: 1202\/1309;\" \/><\/figure>\n\n\n\n<h4 class=\"wp-block-heading\">1. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-8B-Instruct-2512\">Ministral 3 8B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Released in December 2025, this immediately became the model to beat at this size. It\u2019s Apache 2.0 licensed, multimodal (can process images along with text), and optimized for edge deployment. Mistral trained it alongside their larger models, which you will notice in the output quality.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> Ministral is the efficiency pick; its tendency toward shorter, more precise answers makes it the fastest general-purpose model in this class.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">2. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-8B\">Qwen3 8B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">From Alibaba, this model ships with a genuinely useful trick: hybrid thinking modes. You can instruct it to think through complex problems step-by-step or disable reasoning for quick responses. It has a 32K context window natively (extendable to 131K with YaRN, per Qwen\u2019s model card) and strong tool-calling and MCP (Model Context Protocol) support.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> The most versatile 8B model available, specifically optimized for <a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/news\/announcements\/how-we-built-an-ai-powered-business-plan-generator-using-langgraph-langchain\/\">agentic workflows<\/a> where the AI needs to handle complex tools or external data.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">3. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/meta-llama\/Llama-3.1-8B-Instruct\">Llama 3.1 8B Instruct<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">This remains the ecosystem default. Every framework supports it, and every tutorial uses it as an example. However, note the license: Meta\u2019s community agreement is not open-source, and strict usage terms apply.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> The safest bet for compatibility with tutorials and tools, provided you have read the Community License and confirmed your use case complies.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">4. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-Coder-30B-A3B-Instruct\">Qwen3-Coder 30B (MoE)<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">This model exists for just one purpose: <a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/blog\/best-online-resources-learn-to-code\/\">writing code<\/a>. It\u2019s the current generation of Qwen\u2019s dedicated coder line (released mid-2025 under Apache 2.0), a Mixture of Experts design with 30.5 billion total parameters but only 3.3 billion active per token. At 4-bit quantization the weights alone take roughly 15GB (30.5B \u00d7 half a byte), which makes 16GB a tight fit once runtime overhead and KV cache claim their share; pick a smaller quant or offload a few experts if you run long contexts. On 12GB, offloading experts to system RAM works too \u2014 generation slows down, but with only 3.3 billion parameters firing per token, it stays usable.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> The current pick for a local pair programmer; use this if you want Copilot-like suggestions without sending proprietary code to the cloud.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Best self-hosted AI models for 16GB VRAM<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Moving to 16GB allows you to run models that offer a genuine inflection point in reasoning. These models go beyond simple chat to solve problems.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" width=\"1189\" height=\"1323\" data-src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image.png.webp\" alt=\"Grid of four AI model cards for 16GB VRAM \u2014 Ministral 3 14B, Microsoft Phi-4 14B, OpenAI gpt-oss-20b, and Gemma 4 12B\" class=\"wp-image-86679 lazyload\" data-srcset=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image.png.webp 1189w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-270x300.png 270w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-920x1024.png 920w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-768x855.png 768w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-600x668.png.webp 600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-730x812.png.webp 730w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-784x872.png.webp 784w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-16gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-1-image-877x976.png.webp 877w\" data-sizes=\"(max-width: 1189px) 100vw, 1189px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1189px; --smush-placeholder-aspect-ratio: 1189\/1323;\" \/><\/figure>\n\n\n\n<h4 class=\"wp-block-heading\">5. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/mistralai\/Ministral-3-14B-Reasoning-2512\">Ministral 3 14B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">This scales up the architecture of the 8B version with the same focus on efficiency. It offers a 256K context window and a reasoning variant that hits 85% on AIME 2025 (a competition math benchmark).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> A genuine reliability upgrade over the 8B class; the extra VRAM cost pays off significantly in reduced hallucinations and better instruction following.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">6. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/microsoft\/phi-4\">Microsoft Phi-4 14B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Phi-4 ships under the MIT license, the most permissive option available. No usage restrictions whatsoever; it offers strong performance on reasoning tasks and boasts Microsoft\u2019s backing for long-term support.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> The legally safest choice; choose this model if your primary concern is an unrestrictive license for commercial deployment.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">7. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/openai\/gpt-oss-20b\">OpenAI gpt-oss-20b<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">With gpt-oss (August 2025), <a target=\"_blank\" href=\"https:\/\/openai.com\/index\/introducing-gpt-oss\/\">OpenAI released<\/a> its first open-weight language model since GPT-2 in 2019, under an Apache 2.0 license. It uses a <a target=\"_blank\" href=\"https:\/\/huggingface.co\/blog\/moe\">Mixture of Experts (MoE) architecture<\/a>, meaning it has 21 billion parameters but only uses 3.6 billion active parameters per token.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> A clever piece of engineering that delivers the best balance of reasoning capability and inference speed in the 16GB tier.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">8. <a target=\"_blank\" href=\"https:\/\/deepmind.google\/models\/gemma\/gemma-4\/\">Gemma 4 12B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">Google\u2019s Gemma 4 generation (April 2026) finally ships under Apache 2.0, and the 12B \u201cUnified\u201d variant is the multimodal pick for this tier: Google pitches the family\u2019s audio and visual understanding, native function calling, and support for 140 languages, with the 12B\u201331B sizes optimized for consumer GPUs.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">(A note on Llama 4 Scout, which you\u2019ll often see recommended here: despite the \u201c17B\u201d in its name, that\u2019s <em>active<\/em> parameters. The April 2025 release totals 109 billion parameters, and Meta says it fits \u201ca single H100 GPU\u201d \u2014 an 80GB server card rather than a 16GB consumer one.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> The polished option for local multimodal tasks, allowing you to process documents, receipts, and screenshots securely on your own hardware, under a license you can ship with.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Best self-hosted AI models for 24GB+ VRAM<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">If you have an RTX 3090 or 4090, you enter the \u201cPower User\u201d tier, where you can run models that approach frontier-class performance.<\/p>\n\n\n\n<figure class=\"wp-block-image\"><img decoding=\"async\" width=\"1558\" height=\"1010\" data-src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image.png.webp\" alt=\"Comparison of two AI model cards for 24GB+ VRAM \u2014 Qwen3 VL 32B and Gemma 4 26B, both Apache 2.0 licensed\" class=\"wp-image-86680 lazyload\" data-srcset=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image.png.webp 1558w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-300x194.png 300w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-1024x664.png 1024w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-768x498.png 768w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-1536x996.png 1536w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-600x389.png.webp 600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-1200x778.png.webp 1200w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-730x473.png.webp 730w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-1460x946.png.webp 1460w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-784x508.png.webp 784w, https:\/\/www.dreamhost.com\/blog\/wp-content\/smush-webp\/2026\/08\/ai-models-24gb-vram-comparison-dreamops-5042e2a9ecb84acc8e05cd5d1c187355-p3-imageaudit-update-2-image-877x569.png.webp 877w\" data-sizes=\"(max-width: 1558px) 100vw, 1558px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1558px; --smush-placeholder-aspect-ratio: 1558\/1010;\" \/><\/figure>\n\n\n\n<h4 class=\"wp-block-heading\">9. <a target=\"_blank\" href=\"https:\/\/huggingface.co\/Qwen\/Qwen3-VL-32B-Instruct\">Qwen3 VL 32B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">This model targets the 24GB sweet spot specifically. It offers almost everything you\u2019d need: Apache 2.0 licensed, a 256K native context window (per Qwen\u2019s model card), and vision and language in a single model.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> A strong 24GB single-GPU pick for vision and language work, though you\u2019ll need a 4-bit build and conservative context settings to make it fit.<\/p>\n\n\n\n<h4 class=\"wp-block-heading\">10. <a target=\"_blank\" href=\"https:\/\/deepmind.google\/models\/gemma\/gemma-4\/\">Gemma 4 26B<\/a><\/h4>\n\n\n\n<p class=\"wp-block-paragraph\">The larger Gemma 4 sizes are Google\u2019s current answer to \u201cfrontier intelligence on your own hardware.\u201d Google says quantized versions of the 26B and 31B models run natively on consumer GPUs, and its published benchmarks show the 26B beating the previous-generation Gemma 3 27B across the board. The family is multimodal (featuring audio and visual understanding plus native function calling) and, unlike every Gemma generation before it, Apache 2.0 licensed (as of the April 2, 2026 release).<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\u2705Verdict:<\/strong> A high-performance generalist for researchers, hobbyists, and, now that the license is Apache 2.0, commercial products too.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Bonus: distilled reasoning models<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We <em>have<\/em> to mention models like <a target=\"_blank\" href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1-Distill-Qwen-7B\">DeepSeek R1 Distill<\/a>. These exist at multiple sizes and are derived from larger parent models to \u201cthink\u201d (spend more tokens processing) before answering.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Such models are perfect for specific math or logic tasks where accuracy matters more than latency. Licensing is friendlier than you might expect: DeepSeek licenses the R1 series, including distills, under the MIT License, per its model card, though the distills themselves are built on Qwen and Llama bases.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Always read the specific model card before downloading to confirm you are compliant.<\/strong><\/p>\n\n\n\n<h2 id=\"h2_what-tools-should-you-use-to-deploy-local-models\" class=\"wp-block-heading\">What tools should you use to deploy local models?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">You have the hardware and the model. Here are three practical ways to run it:<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">1. <a target=\"_blank\" href=\"https:\/\/ollama.com\/\">Ollama<\/a><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Ollama is widely considered the standard for \u201cgetting it running tonight.\u201d It bundles the engine and model management into a single binary.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>How it works:<\/strong> You install it, type <strong>ollama run llama3 <\/strong>or another model name from <a target=\"_blank\" href=\"https:\/\/ollama.com\/library\">the library<\/a>, and you\u2019re chatting in seconds (depending on the model size and your VRAM).<\/li>\n\n\n\n<li><strong>The killer feature:<\/strong> Simplicity \u2014 it abstracts away all the quantization details and file paths, making it the perfect starting point for beginners.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">2. <a target=\"_blank\" href=\"https:\/\/lmstudio.ai\/\">LM Studio<\/a><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">LM Studio provides a GUI for people who prefer not to live in terminals. You can visualize your model library and manage configurations without memorizing <a target=\"_blank\" href=\"https:\/\/help.dreamhost.com\/hc\/en-us\/sections\/203272488-Command-line-troubleshooting-tools\">command-line<\/a> arguments.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>How it works:<\/strong> You can search for models, download them, configure quantization settings, and run a local API server with a few clicks.<\/li>\n\n\n\n<li><strong>The killer feature:<\/strong> Automatic hardware offloading; it handles integrated GPUs surprisingly well. If you are on a laptop with a modest dedicated GPU or Apple Silicon, LM Studio detects your hardware and automatically splits the model between your CPU and GPU.<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">3. <a target=\"_blank\" href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">llama.cpp Server<\/a><\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">If you want the raw power of open-source without any \u201c<a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/blog\/what-is-the-open-web\/\">walled garden<\/a>,\u201d you can run llama.cpp directly using its built-in server mode. This is often preferred by power users because it eliminates the middleman.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>How it works:<\/strong> You download the llama-server binary, point it at your model file, and it spins up a local web server that\u2019s lightweight and has zero unnecessary dependencies.<\/li>\n\n\n\n<li><strong>The killer feature:<\/strong> Native OpenAI compatibility; with a simple command, you instantly get an OpenAI-compatible API endpoint. You can plug this directly into dictation apps, VS Code extensions, or any tool built for ChatGPT, and it just works.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"h2_how-do-you-build-a-complete-self-hosted-ai-stack\" class=\"wp-block-heading\">How do you build a complete self-hosted AI stack?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A model runner is just the engine. A complete self-hosted AI stack adds a vector database for retrieval and a workflow layer for automation, usually wired together with Docker Compose.<\/strong> That\u2019s the difference between \u201cI can chat with a model\u201d and \u201cmy documents, automations, and agents all run privately on my hardware.\u201d<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Three layers cover most private RAG and agent setups:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Runtime:<\/strong> Ollama or llama.cpp serving your model behind an OpenAI-compatible API (covered above).<\/li>\n\n\n\n<li><strong>Vector database:<\/strong> An open-source store like Qdrant holds embeddings of your documents so the model can search and cite them, which forms the backbone of private RAG.<\/li>\n\n\n\n<li><strong>Workflow layer:<\/strong> A low-code tool like n8n (plus PostgreSQL for state) turns the pieces into pipelines. One practical pattern: a watched folder that embeds every new document into Qdrant, so the model can answer questions about a contract minutes after you save it.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">You don\u2019t have to assemble this by hand. n8n\u2019s <a target=\"_blank\" href=\"https:\/\/github.com\/n8n-io\/self-hosted-ai-starter-kit\">Self-hosted AI Starter Kit<\/a> is an open-source Docker Compose template that bundles n8n, Ollama, Qdrant, and PostgreSQL in one command, with install paths for NVIDIA GPUs, AMD GPUs on Linux, Apple Silicon, and plain CPUs.<\/p>\n\n\n\n<h2 id=\"h2_when-should-you-move-from-local-hardware-to-cloud-infrastructure\" class=\"wp-block-heading\">When should you move from local hardware to cloud infrastructure?<\/h2>\n\n\n\n<figure class=\"wp-block-image size-full\"><img decoding=\"async\" width=\"1600\" height=\"821\" data-src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x.webp\" alt=\"Funnel graphic comparing cloud vs local AI: team\/server use on left, solo\/local privacy on right.\" class=\"wp-image-79443 lazyload\" data-srcset=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x.webp 1600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-300x154.webp 300w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-1024x525.webp 1024w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-768x394.webp 768w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-1536x788.webp 1536w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-600x308.webp 600w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-1200x616.webp 1200w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-730x375.webp 730w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-1460x749.webp 1460w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-784x402.webp 784w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-1568x805.webp 1568w, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/05-Stay-Local-or-Move-to-the-Cloud__1x-877x450.webp 877w\" data-sizes=\"(max-width: 1600px) 100vw, 1600px\" src=\"data:image\/svg+xml;base64,PHN2ZyB3aWR0aD0iMSIgaGVpZ2h0PSIxIiB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9zdmciPjwvc3ZnPg==\" style=\"--smush-placeholder-width: 1600px; --smush-placeholder-aspect-ratio: 1600\/821;\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">Local deployment has limits, and knowing them saves you time and money.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Single-user workloads run great locally, because it\u2019s you and your laptop against the world. Privacy\u2019s absolute, latency\u2019s low, and your cost is zero after hardware. However, multi-user scenarios get complicated fast.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>\ud83d\udcb0Reality check:<\/strong> do the break-even math before buying a GPU for savings alone. Hosted inference is cheap at low volume \u2014 Mistral prices Ministral 3 14B at $0.20 per million tokens on its API (Mistral model card, December 2025), so even a heavy solo user burning 1 million tokens a day spends about $6 a month at per-token rates. A multi-hundred-dollar GPU takes years to pay for itself against that. The math flips when you\u2019re replacing a chatbot subscription (typically around $20\/month), running sustained daily workloads, or when privacy is the actual point. Electricity stays a small line item: a GPU drawing 350W for 2 hours a day adds about 21 kWh a month.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Two people querying the same model might work; 10 people will not. GPU memory doesn\u2019t multiply when you add users. Concurrent requests queue up, latency spikes, and everyone gets frustrated. Furthermore, long context plus speed creates impossible tradeoffs. KV cache scales linearly with context length \u2014 processing 100K tokens of context eats VRAM that could be running inference.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>If you need to build a production service, the tooling changes:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>vLLM:<\/strong> Provides high-throughput inference with OpenAI-compatible APIs, production-grade serving, and optimizations consumer tools skip (like PagedAttention).<\/li>\n\n\n\n<li><strong>SGLang:<\/strong> Focuses on structured generation and constrained outputs, essential for applications that must output valid JSON.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">These tools expect server-grade infrastructure, and production inference generally needs a GPU-equipped server, so confirm the GPU spec before you rent anything. DreamHost\u2019s <a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/hosting\/dedicated\/\">dedicated servers<\/a> are built around CPU, RAM, and SSD configurations, which suits the orchestration layer (n8n, Qdrant, PostgreSQL) and smaller CPU-friendly quantized models. Either way, renting a server beats exposing your home network to the internet.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Here\u2019s a quick way to decide:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Run local:<\/strong> If your goal is one user, privacy, and learning.<\/li>\n\n\n\n<li><strong>Rent infrastructure:<\/strong> If your goal is a service + concurrency + reliability.<\/li>\n<\/ul>\n\n\n\n<h2 id=\"h2_start-building-your-self-hosted-llm-lab-today\" class=\"wp-block-heading\">Start building your self-hosted LLM lab today<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">You run models at home because you want zero latency, zero API bills, and total data privacy. But your GPU becomes the physical boundary. So, if you try to force a 32B model into 12GB of VRAM, your system will crawl or crash.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Instead, use your local machine to prototype, fine-tune your prompts, and vet model behavior.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Once you need to share that model with a team or guarantee it stays online while you sleep, stop fighting your hardware and move the workload to a <a target=\"_blank\" href=\"https:\/\/www.dreamhost.com\/hosting\/dedicated\/\">dedicated server<\/a> designed for 24\/7 uptime.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">You keep the control that made local hosting appealing: DreamHost Dedicated hosting comes with full root and shell access, so you decide what runs on the server and what gets logged. And you also skip the upfront hardware costs and setup.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Here are your next steps:<\/strong><\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Audit your VRAM:<\/strong> Open your task manager or run nvidia-smi (on a Mac, check your unified memory under About This Mac). That number determines your model list. Everything else is secondary.<\/li>\n\n\n\n<li><strong>Test a 7B model:<\/strong> Download Ollama or LM Studio. Run Qwen3 or Ministral at 4-bit quantization to establish your performance baseline.<\/li>\n\n\n\n<li><strong>Identify your bottleneck:<\/strong> If your context windows are hitting memory limits or your fan sounds like a jet engine, evaluate if you\u2019ve outgrown local hosting. High-concurrency tasks belong on dedicated servers, and you may just need to make the switch.<\/li>\n<\/ul>\n\n\n\n\n<div class=\"article-cta-shared article-cta-small article-cta--product\">\n\t<div class=\"tr-img-wrap-outer jsLoading\"><img decoding=\"async\" class=\"js-img-lazy \" src=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/themes\/blog2018\/assets\/img\/lazy-loading-transparent.webp\" data-srcset=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/03\/product-cta-dedicated-hosting-877x586.webp 1x, https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/03\/product-cta-dedicated-hosting.webp 2x\"  \/><\/div>\n\n\t<a href='https:\/\/www.dreamhost.com\/hosting\/dedicated\/' class='link-top' target='_blank' rel='noopener noreferrer'>\n\t\t<span>Dedicated Hosting<\/span>\n\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" viewBox=\"0 0 384 512\" width=\"15\"><path d=\"M342.6 233.4c12.5 12.5 12.5 32.8 0 45.3l-192 192c-12.5 12.5-32.8 12.5-45.3 0s-12.5-32.8 0-45.3L274.7 256 105.4 86.6c-12.5-12.5-12.5-32.8 0-45.3s32.8-12.5 45.3 0l192 192z\"\/><\/svg>\n\t<\/a>\n\n\t<div class=\"content-btm\">\n\t\t<h2 class=\"h2--md\">\n\t\t\tUltimate in Power, Security, and Control\n\t\t<\/h2>\n\t\t<p class=\"p--md\">\n\t\t\tDedicated servers from DreamHost use the best hardware\r\nand software available to ensure your site is always up, and always fast.\n\t\t<\/p>\n\n\t\t        <a\n            href=\"https:\/\/www.dreamhost.com\/hosting\/dedicated\/\"\n                        class=\"btn btn--white-outline btn--sm btn--round\"\n                                    target=\"_blank\"\n            rel=\"noopener noreferrer\"\n            >\n                            See More                    <\/a>\n\n\t<\/div>\n<\/div>\n\n\n<h2 id=\"h2_frequently-asked-questions-about-self-hosted-ai-models\" class=\"wp-block-heading\">Frequently asked questions about self-hosted AI models<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Can I run an LLM on 8GB VRAM?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Yes. Qwen3 4B, Ministral 3B, and other sub-7B models run comfortably. Quantize to Q4 and keep context windows reasonable. Performance won\u2019t match larger models, but functional local AI is absolutely possible on entry-level GPUs.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What model should I use for 12GB?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Ministral 8B is the efficiency winner. And if you\u2019re doing heavy agentic work or tool-use, Qwen3 8B has strong tool-calling and Model Context Protocol (MCP) support, making it the go-to in this weight class.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">What is the best local AI model?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For most people, Ministral 3 8B: Apache 2.0 licensed, multimodal, and fast on a 12GB GPU. Coders should look at Qwen3-Coder 30B (MoE), and anyone with a 24GB+ card should try Qwen3 VL 32B for near-frontier quality.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">When should I use hosted inference instead of local?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">When the model doesn\u2019t fit in your VRAM, even when quantized \u2014 when you need to serve multiple concurrent users, when context requirements exceed what your GPU can handle, or when you need service-grade reliability with SLOs and support.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Run AI models locally on your GPU. Discover the best self-hosted LLMs for 8GB, 12GB, 16GB, and 24GB+ VRAM, plus when to graduate to real infrastructure.<\/p>\n","protected":false},"author":1006,"featured_media":79438,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_yoast_wpseo_opengraph-title":"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost","_yoast_wpseo_opengraph-description":"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.","_yoast_wpseo_twitter-title":"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost","_yoast_wpseo_twitter-description":"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.","toc_headlines":"[[\"h2_what-is-self-hosted-ai\",\"What is self-hosted AI?\"],[\"h-what-s-the-difference-between-open-source-open-weights-and-terms-based-ai\",\"What\u2019s the difference between open-source, open-weights, and terms-based AI?\"],[\"h2_how-does-gpu-memory-determine-which-models-you-can-run\",\"How does GPU memory determine which models you can run?\"],[\"h2_what-can-you-expect-from-different-vram-configurations\",\"What can you expect from different VRAM configurations?\"],[\"h2_the-10-best-self-hosted-ai-models-you-can-run-at-home\",\"The 10 best self-hosted AI models you can run locally\"],[\"h2_what-tools-should-you-use-to-deploy-local-models\",\"What tools should you use to deploy local models?\"],[\"h2_how-do-you-build-a-complete-self-hosted-ai-stack\",\"How do you build a complete self-hosted AI stack?\"],[\"h2_when-should-you-move-from-local-hardware-to-cloud-infrastructure\",\"When should you move from local hardware to cloud infrastructure?\"],[\"h2_start-building-your-self-hosted-llm-lab-today\",\"Start building your self-hosted LLM lab today\"],[\"h2_frequently-asked-questions-about-self-hosted-ai-models\",\"Frequently asked questions about self-hosted AI models\"]]","hide_toc":false,"show_updated_at":"1","footnotes":""},"categories":[14839],"tags":[],"class_list":["post-42964","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.4 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost<\/title>\n<meta name=\"description\" content=\"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost\" \/>\n<meta property=\"og:description\" content=\"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/\" \/>\n<meta property=\"og:site_name\" content=\"DreamHost Blog\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/DreamHost\/\" \/>\n<meta property=\"article:published_time\" content=\"2026-02-09T09:00:00+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-08-04T04:48:15+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/1220-x-628-OGIMAGE_The-10-Best-Self-Hosted-AI-Models-You-Can-Run-at-Home.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"628\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"Brian Andrus\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:title\" content=\"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost\" \/>\n<meta name=\"twitter:description\" content=\"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.\" \/>\n<meta name=\"twitter:creator\" content=\"@dreamhost\" \/>\n<meta name=\"twitter:site\" content=\"@dreamhost\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Brian Andrus\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"13 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost","description":"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/","og_locale":"en_US","og_type":"article","og_title":"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost","og_description":"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.","og_url":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/","og_site_name":"DreamHost Blog","article_publisher":"https:\/\/www.facebook.com\/DreamHost\/","article_published_time":"2026-02-09T09:00:00+00:00","article_modified_time":"2026-08-04T04:48:15+00:00","og_image":[{"width":1200,"height":628,"url":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/1220-x-628-OGIMAGE_The-10-Best-Self-Hosted-AI-Models-You-Can-Run-at-Home.webp","type":"image\/webp"}],"author":"Brian Andrus","twitter_card":"summary_large_image","twitter_title":"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost","twitter_description":"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.","twitter_creator":"@dreamhost","twitter_site":"@dreamhost","twitter_misc":{"Written by":"Brian Andrus","Est. reading time":"13 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#article","isPartOf":{"@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/"},"author":{"name":"Brian Andrus","@id":"https:\/\/www.dreamhost.com\/blog\/#\/schema\/person\/bea7bcc7f0758793822403e65b57489e"},"headline":"The 10 Best Self-Hosted AI Models You Can Run at Home","datePublished":"2026-02-09T09:00:00+00:00","dateModified":"2026-08-04T04:48:15+00:00","mainEntityOfPage":{"@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/"},"wordCount":3583,"publisher":{"@id":"https:\/\/www.dreamhost.com\/blog\/#organization"},"image":{"@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#primaryimage"},"thumbnailUrl":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/1460-x-1095-BLOG-HERO_The-10-Best-Self-Hosted-AI-Models-You-Can-Run-at-Home.webp","articleSection":["AI"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/","url":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/","name":"Self-Hosted AI: 10 Best Local AI Models to Run - DreamHost","isPartOf":{"@id":"https:\/\/www.dreamhost.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#primaryimage"},"image":{"@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#primaryimage"},"thumbnailUrl":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/1460-x-1095-BLOG-HERO_The-10-Best-Self-Hosted-AI-Models-You-Can-Run-at-Home.webp","datePublished":"2026-02-09T09:00:00+00:00","dateModified":"2026-08-04T04:48:15+00:00","description":"Self-hosted AI lets you run models like Qwen3 and Phi-4 locally \u2014 DreamHost breaks down the VRAM requirements, quantization, and tools you need.","breadcrumb":{"@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.dreamhost.com\/blog\/open-source-ai\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#primaryimage","url":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/1460-x-1095-BLOG-HERO_The-10-Best-Self-Hosted-AI-Models-You-Can-Run-at-Home.webp","contentUrl":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2024\/01\/1460-x-1095-BLOG-HERO_The-10-Best-Self-Hosted-AI-Models-You-Can-Run-at-Home.webp","width":1460,"height":1095,"caption":"The 10 Best Self-Hosted AI Models You Can Run at Home"},{"@type":"BreadcrumbList","@id":"https:\/\/www.dreamhost.com\/blog\/open-source-ai\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.dreamhost.com\/blog\/"},{"@type":"ListItem","position":2,"name":"The 10 Best Self-Hosted AI Models You Can Run at Home"}]},{"@type":"WebSite","@id":"https:\/\/www.dreamhost.com\/blog\/#website","url":"https:\/\/www.dreamhost.com\/blog\/","name":"DreamHost Blog","description":"","publisher":{"@id":"https:\/\/www.dreamhost.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.dreamhost.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.dreamhost.com\/blog\/#organization","name":"DreamHost","url":"https:\/\/www.dreamhost.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.dreamhost.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/dhblog.dream.press\/blog\/wp-content\/uploads\/2019\/01\/dh_logo-blue-2.png","contentUrl":"https:\/\/dhblog.dream.press\/blog\/wp-content\/uploads\/2019\/01\/dh_logo-blue-2.png","width":1200,"height":168,"caption":"DreamHost"},"image":{"@id":"https:\/\/www.dreamhost.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/DreamHost\/","https:\/\/x.com\/dreamhost","https:\/\/www.instagram.com\/dreamhost\/","https:\/\/www.linkedin.com\/company\/dreamhost\/","https:\/\/www.youtube.com\/user\/dreamhostusa"]},{"@type":"Person","@id":"https:\/\/www.dreamhost.com\/blog\/#\/schema\/person\/bea7bcc7f0758793822403e65b57489e","name":"Brian Andrus","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2023\/10\/brian-andrus-150x150.jpg","url":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2023\/10\/brian-andrus-150x150.jpg","contentUrl":"https:\/\/www.dreamhost.com\/blog\/wp-content\/uploads\/2023\/10\/brian-andrus-150x150.jpg","caption":"Brian Andrus"},"description":"Brian is a Cloud Engineer at DreamHost, primarily responsible for cloudy things. In his free time he enjoys navigating fatherhood, cutting firewood, and self-hosting whatever he can.","url":"https:\/\/www.dreamhost.com\/blog\/author\/brianandrus\/"}]}},"lang":"en","translations":{"en":42964,"es":42974,"ru":50747,"de":54697,"uk":54706,"pt":54716,"pl":54749,"it":68118,"fr":69831,"nl":69850},"pll_sync_post":{},"_links":{"self":[{"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/posts\/42964","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/users\/1006"}],"version-history":[{"count":1,"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/posts\/42964\/revisions"}],"predecessor-version":[{"id":86681,"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/posts\/42964\/revisions\/86681"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/media\/79438"}],"wp:attachment":[{"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/media?parent=42964"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/categories?post=42964"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.dreamhost.com\/blog\/wp-json\/wp\/v2\/tags?post=42964"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}