About you
About your app
One last thing
Tell us a bit about yourself and your АІ app below

We will finalize the submission, and our team will follow up within 10 business days. If you want to submit multiple apps, please fill in the “About your app” section separately for each one.

Select Country
Continue
App 1
Remove app
Select Country
Add another app
Continue
Thanks! We’ve got your submission
You’ll hear back from our team within 10 business days.
Go back
Oops! Something went wrong while submitting the form.

Meta Llama is the most widely used open-source LLM in the world. Llama 4 Scout offers a 10-million-token context window — 10x more than anything else available. Maverick delivers frontier-grade general intelligence. The whole family is free to download, fine-tune, and deploy on your own terms.

Context window
131K tokens
Pricing
Per-token
License
Open-weight

Llama models compared

Model
Arena ELO ±CI
Rank
Votes
Context
$ / 1M in
$ / 1M out
License
Llama 3.3 70b Instruct
1317.8 ±3.5
257
54368
131K tokens
0.1
0.32
Open-weight
Llama 3.1 70b Instruct
1293.3 ±3.8
282
55240
131K tokens
0.4
0.4
Open-weight
Llama 3.1 8b Instruct
1211.2 ±4.2
336
49605
131K tokens
0.05
0.08
Open-weight
Data source:
· Data as of
25 Sep 2026
Learn more about the methodology

What Llama is best at

Truly open-source with permissive licensing
Meta Llama ships under a license that allows commercial use, fine-tuning, and redistribution. You get downloadable weights and published training details. The community has already built thousands of specialized variants on top of it.
10-million-token context window (Scout)
Llama AI Scout can process 10 million tokens at once. That's entire document archives, massive codebases, or years of conversation logs — no chunking required. Nothing else, whether open or proprietary, operates at this scale.
The largest open-source ecosystem
As the most popular open-source LLM, Llama has more fine-tuned variants on Hugging Face, wider framework support (Ollama, vLLM, TensorRT-LLM), and more community tooling than any alternative. Picking Llama means picking the ecosystem with the most momentum behind it.

How to use Llama via Setapp AI Gateway

One API, all models
Step 1
Access ChatGPT, Claude, Gemini, DeepSeek, Grok, Llama, Mistral, Qwen, Copilot-class models, and Perplexity through a single endpoint. One key, one integration, every model.
Transparent, unified pricing
Step 2
No surprise bills, no per-model accounts. Setapp Al Gateway consolidates billing across providers with clear per-token pricing and usage dashboards.
Smart routing
Step 3
Automatically route requests to the best model for each task — or choose manually. Optimize for cost, speed, or capability without changing your code.
Zero vendor lock-in
Step 4
Switch between models with a parameter change. Your integration stays the same. When a new model launches, it's available in your existing setup - no migration needed.

What people build with Llama

extract
On-premise and private cloud deployments
Run Llama locally on your hardware — no API calls, keeping data within your network, with no rate limits or usage-based charges. This setup fully removes third-party data risks for regulated industries and government agencies.
ship
Fine-tuned domain-specific models
Start with a Llama base and fine-tune it on your proprietary data — medical records, legal precedents, customer conversations, internal docs. The result is a model that understands your domain natively.
refine
Edge and embedded AI
Smaller Llama variants run on consumer GPUs, phones, and edge devices. That lets you build AI features that work offline, respond instantly, and never send data off-device.
trust
High-volume production pipelines
When you self-host, compute is your only cost. For workloads with millions of daily requests (classification, extraction, summarization), self-hosted Llama has the best unit economics available.

Common questions

What is Llama AI and who develops it?
Llama stands for Large Language Model Meta AI. Meta develops and releases it as open-source. The Llama 4 family includes Scout (10M-token context), Maverick (1M context, general-purpose), and Behemoth (the largest and most capable). It's the most downloaded open-source LLM globally.
Is Meta Llama really free to use commercially?
For the vast majority of businesses, yes. The license permits commercial use, fine-tuning, and deployment. The one exception: companies with over 700 million monthly active users need a separate agreement with Meta.
Open source LLM — why choose Llama over proprietary models?
You get full control over infrastructure, zero per-token fees, and complete data privacy. You can also fine-tune for your specific domain and run offline. The tradeoff is that you handle deployment and maintenance yourself.
What hardware do I need to run Llama?
It depends on the model size. Smaller Llama models can run on a modern desktop or laptop with a capable GPU, while larger models require more powerful hardware or cloud infrastructure. If you're just getting started, tools such as Ollama simplify running smaller Llama models on your local machine.
How does Llama compare to DeepSeek and other open models?
Llama wins on community size, number of fine-tuned variants, and breadth of tooling. DeepSeek often wins on cost-performance at the API level. Qwen leads on multilingual breadth. For teams starting with open-source, Llama's ecosystem usually makes it the safest first choice.

Browse the best Al models

Every large language model available through Setapp - in one chat, under one subscription. Click any model for benchmarks, pricing, sample prompts, and comparisons.

Route to Llama via Setapp AI Gateway

Access hosted Llama models alongside proprietary options. Compare performance, control costs, stay flexible.

Al models for developers - one integration, every provider

Setapp Al Gateway gives your app access to all models listed above through a single OpenAl-compatible endpoint. No separate API keys, 
no provider
contracts, no billing logic to build.

Al models for developers - one integration,
every provider

Setapp Al Gateway gives your app access to all models listed above through a single OpenAl-compatible endpoint. No separate API keys, 
no provider contracts, no billing logic to build.