All models

Qwen3.7 Flash vs Llama 3 8B Lunaris

Qwen3.7 Flash at $0.035 in and $0.149 out and Llama 3 8B Lunaris at $0.046 in and $0.058 out — per million tokens, from the same wallet.

Q

Qwen3.7 Flash

Qwen

0.03¢

for a client proposal

$0.035 in · $0.149 out / 1M

1M context · released 2 weeks ago

Full page →
S

Llama 3 8B Lunaris

Sao10k

0.03¢

for a client proposal

$0.046 in · $0.058 out / 1M

8K context · released August 2024

Full page →

Asking all both at once — identical prompt, identical context — costs about 0.06¢ for a typical client proposal.

At a glance

  • Cheapest input: Qwen3.7 Flash at $0.035 per million tokens.
  • Cheapest output: Llama 3 8B Lunaris at $0.058 per million tokens.
  • Largest context: Qwen3.7 Flash at 1M tokens — about 1,400 pages.

Cost per task

Estimated totals at the rates above. Solid bar is input, lighter bar is output.

Quick code fix

2K in · 500 out
Qwen3.7 Flash
0.01¢
Llama 3 8B Lunaris
0.01¢

Draft a client proposal from notes

5K in · 800 out
Qwen3.7 Flash
0.03¢
Llama 3 8B Lunaris
0.03¢

Context-heavy session (Alyph workspace)

155K in · 4K out
Qwen3.7 Flash
Llama 3 8B Lunaris
needs 155K ctx

Long-context rate applied at this input size.

Analyze a 50-page legal contract

35K in · 1K out
Qwen3.7 Flash
0.45¢
Llama 3 8B Lunaris
needs 35K ctx

Long-context rate applied at this input size.

Large architecture refactor

250K in · 3K out
Qwen3.7 Flash
Llama 3 8B Lunaris
needs 250K ctx

Long-context rate applied at this input size.

Solid part of each bar is input tokens, the lighter part is output. Models that don’t fit a task are marked instead of priced.

Cost vs Token Amount

Drag the slider to adjust the split between input and output workload.

80% Input20% Output
free$0.18$0.370 tokens250k500k750k1000k
Qwen3.7 Flash
Llama 3 8B Lunaris

Adjust the ratio slider to change how output tokens influence the final price for equivalent workloads.

Pricing

Qwen3.7 Flash
Llama 3 8B Lunaris
Input / 1M tokens
$0.035
$0.046
Output / 1M tokens
$0.149
$0.058
Cached input / 1M
$0.0069

Specs

Qwen3.7 Flash
Llama 3 8B Lunaris
Provider
Released
2 weeks ago
August 2024
Context window
1M
8K
Max output
66K
16K
Input
Text, Images, and Video
Text
Output
Text
Text
Reasoning
Shows its thinking
Answers directly
Knowledge cutoff
2023-12-31

Try it with real numbers

2K tokens in · 500 tokens out.

Qwen3.7 Flash

Qwen

0.01¢

for this task

48% input · 52% output

Full pricing →

Llama 3 8B Lunaris

Sao10k

0.01¢

for this task

76% input · 24% output

Full pricing →

All 2, same prompt and context: 0.03¢. That is the entire cost of the comparison.

Your $5 welcome credit covers about 18,903 of these.

Which should you pick?

For long documents, Qwen3.7 Flash has the largest context window here — 1M tokens.

On price, Llama 3 8B Lunaris is the cheapest of the two on a typical task, at about 0.03¢.

For images or files, Qwen3.7 Flash can read them directly; Llama 3 8B Lunaris is text-only.

If you want to watch the model think, Qwen3.7 Flash shows reasoning; the other answers directly.

Or don’t pick. Send the identical prompt to all both, read the answers side by side, and keep the winner. How to run a fair bake-off →

Questions, answered

Which is cheaper, Qwen3.7 Flash or Llama 3 8B Lunaris?

On a typical client proposal (5K tokens in, 800 out): Llama 3 8B Lunaris at 0.03¢; Qwen3.7 Flash at 0.03¢. Long prompts can change the order when long-context rates apply.

Which has the largest context window?

Qwen3.7 Flash — 1M tokens, about 1,400 pages.

Can I run Qwen3.7 Flash and Llama 3 8B Lunaris side by side on Alyph?

Yes — that is what Alyph is built for: the identical prompt with identical context to all 2, answers rendered side by side, billed from one wallet. This exact combination costs about 0.06¢ for a typical task.

Ask both at once

Identical prompt, identical context, answers side by side — about 0.06¢ for a typical client proposal. Free to start, with $5 of credit.