model family

Llama 3.2

Published by meta-llama. 8 members have been measured on real machines. Every number below was signed by the machine that produced it — none of it is from a model card.

what it takes, what it returns
text → text

Declared by the network’s catalog, not measured — a signature says what this family would be asked. Whether it answered is the signed capability flags on each member.

8 of 39 builds measured · 31 still open
what each build proved
1b · Q4_K_M · bartowski · Tool calling — proved1b · Q4_K_M · bartowski · Tool loop finished — proved1b · Q4_K_M · bartowski · Faithful to tool results — proved1b · Q4_K_M · bartowski · Structured output — proved1b · Q4_K_M · bartowski · Sustained output — proved1b · Q4_K_M · bartowski · Writes code — proved1b · Q4_K_M · bartowski · Synthesises sources — proved1b · Q4_K_M · bartowski · Follows instructions — proved1b · Q8_0 · Tool calling — proved1b · Q8_0 · Tool loop finished — proved1b · Q8_0 · Faithful to tool results — proved1b · Q8_0 · Structured output — proved1b · Q8_0 · Sustained output — proved1b · Q8_0 · Writes code — proved1b · Q8_0 · Synthesises sources — proved1b · Q8_0 · Follows instructions — proved3b · IQ3_XS · Tool calling — proved3b · IQ3_XS · Tool loop finished — proved3b · IQ3_XS · Faithful to tool results — asked, failed3b · IQ3_XS · Structured output — proved3b · IQ3_XS · Sustained output — proved3b · IQ3_XS · Writes code — proved3b · IQ3_XS · Synthesises sources — asked, failed3b · IQ3_XS · Follows instructions — proved1b · IQ3_XS · Tool calling — proved1b · IQ3_XS · Tool loop finished — proved1b · IQ3_XS · Faithful to tool results — proved1b · IQ3_XS · Structured output — proved1b · IQ3_XS · Sustained output — proved1b · IQ3_XS · Writes code — proved1b · IQ3_XS · Synthesises sources — asked, failed1b · IQ3_XS · Follows instructions — proved1b · IQ1_S · Tool calling — asked, failed1b · IQ1_S · Tool loop finished — never asked1b · IQ1_S · Faithful to tool results — never asked1b · IQ1_S · Structured output — asked, failed1b · IQ1_S · Sustained output — asked, failed1b · IQ1_S · Writes code — asked, failed1b · IQ1_S · Synthesises sources — asked, failed1b · IQ1_S · Follows instructions — asked, failed1b · Q4_K_M · MaziyarPanahi · Tool calling — proved1b · Q4_K_M · MaziyarPanahi · Tool loop finished — proved1b · Q4_K_M · MaziyarPanahi · Faithful to tool results — proved1b · Q4_K_M · MaziyarPanahi · Structured output — proved1b · Q4_K_M · MaziyarPanahi · Sustained output — proved1b · Q4_K_M · MaziyarPanahi · Writes code — proved1b · Q4_K_M · MaziyarPanahi · Synthesises sources — asked, failed1b · Q4_K_M · MaziyarPanahi · Follows instructions — proved1b · F16 · Tool calling — proved1b · F16 · Tool loop finished — proved1b · F16 · Faithful to tool results — proved1b · F16 · Structured output — proved1b · F16 · Sustained output — proved1b · F16 · Writes code — proved1b · F16 · Synthesises sources — asked, failed1b · F16 · Follows instructions — proved3b · IQ1_S · Tool calling — asked, failed3b · IQ1_S · Tool loop finished — never asked3b · IQ1_S · Faithful to tool results — never asked3b · IQ1_S · Structured output — proved3b · IQ1_S · Sustained output — proved3b · IQ1_S · Writes code — asked, failed3b · IQ1_S · Synthesises sources — asked, failed3b · IQ1_S · Follows instructions — asked, failed
Tool calling
Tool loop finished
Faithful to tool results
Structured output
Sustained output
Writes code
Synthesises sources
Follows instructions
hf.co/bartowski/Llama-3.2-1B-Instruct-GGUF:Q4_K_M1b · Q4_K_M · 3 machines · 1.24B
16,384recalled
served
103.9best tok/s
103.9 tok/s · unnamed machine11.6 tok/s · unnamed machine4.7 tok/s · unnamed machine
4.7103.9 tok/s across 3 machines · ×22.0 spreadunnamed machine
recalled 16k⌐ declared 128k · unsignedadvertises 8× what it proved
llama3.2:1b1b · Q8_0 · 5 machines · 1.2B
8,192recalled
served
49.9best tok/s
18.6 tok/s · unnamed machine13.1 tok/s · x86 · CPU · Hyderabad9.7 tok/s · unnamed machine4.1 tok/s · unnamed machine1.8 tok/s · unnamed machine
1.818.6 tok/s across 5 machines · ×10.5 spreadunnamed machinex86 · CPU
recalled 8k⌐ declared 128k · unsignedadvertises 16× what it proved
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ3_XS1b · IQ3_XS · 3 machines · 1.24B
4,096recalled
served
107.5best tok/s
107.5 tok/s · unnamed machine17.7 tok/s · unnamed machine1.5 tok/s · unnamed machine
1.5107.5 tok/s across 3 machines · ×72.9 spreadunnamed machine
recalled 4k⌐ declared 128k · unsignedadvertises 32× what it proved
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ1_S1b · IQ1_S · 3 machines · 1.24B
recalled
served
105.0best tok/s
105.0 tok/s · unnamed machine10.7 tok/s · unnamed machine3.3 tok/s · unnamed machine
3.3105.0 tok/s across 3 machines · ×32.2 spreadunnamed machine
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q4_K_M1b · Q4_K_M · 2 machines · 1.24B
recalled
served
113.3best tok/s
113.3 tok/s · unnamed machine8.9 tok/s · unnamed machine
8.9113.3 tok/s across 2 machines · ×12.8 spreadunnamed machine

Builds this network knows of

Declared, not measured — this is what the catalog says exists and what it would cost a machine to hold. Nothing here is signed, and none of it claims the build works.

how much memory does your machine have

38 of 39 builds fit a 16 GB machine — budgeting 70% of it, or 11.2 GiB, because the cache and the runtime want the rest.

  • hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ1_SIQ1_Sollamameasured herequantized by MaziyarPanahi
    0.4 GiB
  • hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ3_XSIQ3_XSollamameasured herequantized by MaziyarPanahi
    0.6 GiB
  • mlx-community/Llama-3.2-1B-Instruct-4bit4BITmlxderived by mlx-community
    0.6 GiB
  • hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q3_K_LQ3_K_Lollamaquantized by MaziyarPanahi
    0.7 GiB
  • hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q4_K_SQ4_K_Sollamaquantized by MaziyarPanahi
    0.7 GiB
  • hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q4_K_MQ4_K_Mollamameasured herequantized by MaziyarPanahi
    0.8 GiB
  • hf.co/hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF:Q4_K_MQ4_K_Mollamaderived by hugging-quants
    0.8 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:IQ1_SIQ1_Sollamaquantized by unsloth
    0.8 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:IQ2_XXSIQ2_XXSollamaquantized by unsloth
    1.0 GiB
  • hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q8_0Q8_0ollamaquantized by MaziyarPanahi
    1.2 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:IQ3_XXSIQ3_XXSollamaquantized by unsloth
    1.3 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q2_KQ2_Kollamaderived by bartowski
    1.4 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:IQ3_XSIQ3_XSollamameasured herederived by bartowski
    1.5 GiB
  • mlx-community/Llama-3.2-3B-Instruct-4bit4BITmlxderived by mlx-community
    1.7 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_0Q4_0ollamaquantized by unsloth
    1.8 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q3_K_LQ3_K_Lollamaderived by bartowski
    1.8 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_MQ4_K_Mollamaquantized by unsloth
    1.9 GiB
  • hf.co/hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF:Q4_K_MQ4_K_Mollamaderived by hugging-quants
    1.9 GiB
  • RedHatAI/Llama-3.2-1B-Instruct-FP8-dynamicFP8vllmquantized by RedHatAI
    1.9 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_0_8_8Q4_0_8_8ollamaderived by bartowski
    2.0 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_0_4_4Q4_0_4_4ollamaderived by bartowski
    2.0 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_K_MQ4_K_Mollamaderived by bartowski
    2.1 GiB
  • unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bitBITSANDBYTESvllmderived by unsloth
    2.2 GiB
  • mlx-community/Llama-3.2-1B-Instruct-bf16BF16mlxderived by mlx-community
    2.3 GiB
  • unsloth/Llama-3.2-1B-InstructBF16vllmderived by unsloth
    2.3 GiB
  • hf.co/bartowski/Llama-3.2-1B-Instruct-GGUF:F16F16ollamameasured herequantized by bartowski
    2.3 GiB
  • mlx-community/Llama-3.2-3B-Instruct-uncensored-6bit6BITmlxderived by mlx-community
    2.4 GiB
  • mlx-community/Llama-3.2-3B-Instruct-8bitBF16mlxderived by mlx-community
    3.2 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:Q8_0Q8_0ollamaquantized by unsloth
    3.2 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q8_0Q8_0ollamaderived by bartowski
    3.6 GiB
  • RedHatAI/Llama-3.2-3B-Instruct-FP8-dynamicFP8vllmquantized by RedHatAI
    4.1 GiB
  • hf.co/leafspark/Llama-3.2-11B-Vision-Instruct-GGUF:Q4_K_MQ4_K_Mollamaquantized by leafspark
    5.6 GiB
  • unsloth/Llama-3.2-3B-InstructBF16vllmderived by unsloth
    6.0 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:F16F16ollamaquantized by unsloth
    6.0 GiB
  • hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:BF16BF16ollamaquantized by unsloth
    6.0 GiB
  • hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:F16F16ollamaderived by bartowski
    6.7 GiB
  • hf.co/leafspark/Llama-3.2-11B-Vision-Instruct-GGUF:Q8_0Q8_0ollamaquantized by leafspark
    9.7 GiB
  • mlx-community/Llama-3.2-11B-Vision-Instruct-8bitBF16mlxderived by mlx-community
    10.6 GiB
  • hf.co/leafspark/Llama-3.2-11B-Vision-Instruct-GGUF:F16F16ollamaquantized by leafspark
    18.2 GiB

Sizes are what the catalog declares, not what a machine measured — fitting on disk is necessary and not sufficient. Nobody here has run these on your machine, and this narrows the candidates rather than certifying one.

the same builds, by member
Llama 3.2 11B
hf.co/leafspark/Llama-3.2-11B-Vision-Instruct-GGUF:Q4_K_MQ4_K_Mollama5.6 GiB on diskquantized by leafspark
hf.co/leafspark/Llama-3.2-11B-Vision-Instruct-GGUF:Q8_0Q8_0ollama9.7 GiB on diskquantized by leafspark
hf.co/leafspark/Llama-3.2-11B-Vision-Instruct-GGUF:F16F16ollama18.2 GiB on diskquantized by leafspark
Llama 3.2 11B · mlx-community
mlx-community/Llama-3.2-11B-Vision-Instruct-8bitBF16mlx10.6 GiB on diskderived by mlx-community
Llama 3.2 1B1.2B
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ1_SIQ1_Sollama375 MiB on diskquantized by MaziyarPanahi
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:IQ3_XSIQ3_XSollama592 MiB on diskquantized by MaziyarPanahi
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q3_K_LQ3_K_Lollama698 MiB on diskquantized by MaziyarPanahi
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q4_K_SQ4_K_Sollama739 MiB on diskquantized by MaziyarPanahi
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q4_K_MQ4_K_Mollama770 MiB on diskquantized by MaziyarPanahi
hf.co/MaziyarPanahi/Llama-3.2-1B-Instruct-GGUF:Q8_0Q8_0ollama1.2 GiB on diskquantized by MaziyarPanahi
RedHatAI/Llama-3.2-1B-Instruct-FP8-dynamicFP8vllm1.9 GiB on diskquantized by RedHatAI
hf.co/bartowski/Llama-3.2-1B-Instruct-GGUF:F16F16ollama2.3 GiB on diskquantized by bartowski
Llama 3.2 1B · hugging-quants
hf.co/hugging-quants/Llama-3.2-1B-Instruct-Q4_K_M-GGUF:Q4_K_MQ4_K_Mollama770 MiB on diskderived by hugging-quants
Llama 3.2 1B · mlx-community
mlx-community/Llama-3.2-1B-Instruct-4bit4BITmlx663 MiB on diskderived by mlx-community
mlx-community/Llama-3.2-1B-Instruct-bf16BF16mlx2.3 GiB on diskderived by mlx-community
Llama 3.2 1B · unsloth
unsloth/Llama-3.2-1B-InstructBF16vllm2.3 GiB on diskderived by unsloth
Llama 3.2 3B3.2B
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:IQ1_SIQ1_Sollama870 MiB on diskquantized by unsloth
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:IQ2_XXSIQ2_XXSollama998 MiB on diskquantized by unsloth
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:IQ3_XXSIQ3_XXSollama1.3 GiB on diskquantized by unsloth
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_0Q4_0ollama1.8 GiB on diskquantized by unsloth
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:Q4_K_MQ4_K_Mollama1.9 GiB on diskquantized by unsloth
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:Q8_0Q8_0ollama3.2 GiB on diskquantized by unsloth
RedHatAI/Llama-3.2-3B-Instruct-FP8-dynamicFP8vllm4.1 GiB on diskquantized by RedHatAI
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:F16F16ollama6.0 GiB on diskquantized by unsloth
hf.co/unsloth/Llama-3.2-3B-Instruct-GGUF:BF16BF16ollama6.0 GiB on diskquantized by unsloth
Llama 3.2 3B · bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q2_KQ2_Kollama1.4 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:IQ3_XSIQ3_XSollama1.5 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q3_K_LQ3_K_Lollama1.8 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_0_8_8Q4_0_8_8ollama2.0 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_0_4_4Q4_0_4_4ollama2.0 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_K_MQ4_K_Mollama2.1 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:Q8_0Q8_0ollama3.6 GiB on diskderived by bartowski
hf.co/bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF:F16F16ollama6.7 GiB on diskderived by bartowski
Llama 3.2 3B · hugging-quants
hf.co/hugging-quants/Llama-3.2-3B-Instruct-Q4_K_M-GGUF:Q4_K_MQ4_K_Mollama1.9 GiB on diskderived by hugging-quants
Llama 3.2 3B · mlx-community
mlx-community/Llama-3.2-3B-Instruct-4bit4BITmlx1.7 GiB on diskderived by mlx-community
mlx-community/Llama-3.2-3B-Instruct-uncensored-6bit6BITmlx2.4 GiB on diskderived by mlx-community
mlx-community/Llama-3.2-3B-Instruct-8bitBF16mlx3.2 GiB on diskderived by mlx-community
Llama 3.2 3B · unsloth
unsloth/Llama-3.2-3B-Instruct-unsloth-bnb-4bitBITSANDBYTESvllm2.2 GiB on diskderived by unsloth
unsloth/Llama-3.2-3B-InstructBF16vllm6.0 GiB on diskderived by unsloth
ask your own gram

Your concierge can answer this from the network rather than from this page. The ask carries the operator id, so it resolves to a capability and an engine — not to a name somebody has to recognise.

What has this network measured about the Llama 3.2 family (epn.family.llama-3-2), and can my gram run any of it? See /inferences/llama-3-2

open a gramx with this ask →epn.family.llama-3-2