Saqr
Frontier cybersecurity intelligence, engineered by ZySec.
A 27.8-billion-parameter dense model — threat intelligence, vulnerability analysis and security knowledge, in both thinking and non-thinking modes, from a model small enough to run on a single machine.
- parameters
- 27.8Bparameters
- multimodal
- Text + imagemultimodal
- context
- 262Kcontext
- thinking + non-thinking
- Dual-modethinking + non-thinking
Saqr is a security decision support tool, not a substitute for professional security judgement.
Ahead of every model in this comparison — general frontier models like GPT-5 and GPT-4.1 included, not just the named cybersecurity specialists.
- GPT-5OpenAI94.1
- GPT-4.1OpenAI93.7
- GPT-5-MiniOpenAI93.2
- Foundation-Sec-8B-ReasoningCisco Foundation AI84.3
Ahead of Foundation-Sec-8B-Reasoning’s published 78.2 on general security knowledge.
Saqr measured this evaluation pass, full question-set coverage. Competitor figures are published by their own vendors or the benchmark's own paper.
Benchmarked against
- Foundation-Sec-8B-Reasoning
- MVMinerva (Llama-3.1-8B)
- GPT-4
- GPT-4.1
- GPT-5
- GPT-5-Mini
- GPT-5-Nano
- o3-mini
- GPT-OSS-120B
- GPT-OSS-20B
- Llama-3.3-70B-Instruct
- Llama-3.1-8B-Instruct
- MSPhi-4
- Foundation-Sec-8B-Instruct
- ChatGPT-4
- ChatGPT-3.5
- Gemini-1.5
- Llama 3-70B
- Llama 3-8B
- GLGLM-4-9B
- DeepSeek-V2-Lite
- Mixtral-8x7B
- YIYi-1.5-34B
- Hunyuan-Turbo
Benchmarks
Measured against the models built for cybersecurity
The comparison that matters most is against models built for this domain. Saqr clears the named comparison figure on 9 of 9 gated benchmarks — Foundation-Sec-8B-Reasoning for eight of them, GPT-4’s own published figure for the one Foundation-Sec does not report.
CTIBench-MCQA
78.5
+9.4vs Foundation-Sec-8B-ReasoningMultiple-choice CTI knowledge across threat frameworks and taxonomies
CTIBench-RCM
77.8
+2.5vs Foundation-Sec-8B-ReasoningMaps CVE vulnerability descriptions to their root-cause CWE entry
CTIBench-VSP
90.1
+4.5vs Foundation-Sec-8B-ReasoningPredicts CVSS vulnerability severity from a text description
CTIBench-ATE
72.1
+23.0vs Foundation-Sec-8B-ReasoningExtracts MITRE ATT&CK techniques from threat reports · no-think mode
Ahead of every frontier model tested
SecEval
94.4
+9.6vs Foundation-Sec-8B-ReasoningMultiple-choice cybersecurity knowledge across nine security domains
Ahead of every frontier model tested
CyberMetric-2000
96.2
+11.9vs Foundation-Sec-8B-ReasoningCybersecurity knowledge questions sourced from standards and RFCs
Ahead of every frontier model tested
CyberMetric-10000
91.4
+2.5vs GPT-4The full 10,000-question CyberMetric knowledge set
SecBench
89.7
+17.2vs Foundation-Sec-8B-ReasoningMulti-dimensional cybersecurity knowledge and reasoning questions
MMLU-Security
89.7
+11.5vs Foundation-Sec-8B-ReasoningComputer-security subset of the MMLU knowledge benchmark
Deltas are against each row’s named published figure — Foundation-Sec-8B-Reasoning for eight of the nine, GPT-4’s own CyberMetric-10000 figure for the one it does not report.
Macro average — Saqr across all nine, Foundation-Sec-8B-Reasoning across the eight it reports
Saqr against the security specialist
Foundation-Sec-8B-Reasoning is the reference point that matters most: Cisco’s own security-tuned model. Its published figures are already an average over five sampled trials by its own technical report’s admission, while every Saqr figure here is a single evaluation pass at full question-set coverage.
- Saqr
- Foundation-Sec-8B-Reasoning — published
CTIBench-MCQA
CTIBench-RCM
CTIBench-VSP
CTIBench-ATE
SecEval
CyberMetric-2000
CyberMetric-10000
SecBench
MMLU-Security
View as table
| Benchmark | Saqr | Foundation-Sec-8B-Reasoning — published |
|---|---|---|
| CTIBench-MCQA | 78.5 | 69.1 |
| CTIBench-RCM | 77.8 | 75.3 |
| CTIBench-VSP | 90.1 | 85.6 |
| CTIBench-ATE | 72.1 | 49.1 |
| SecEval | 94.4 | 84.8 |
| CyberMetric-2000 | 96.2 | 84.3 |
| CyberMetric-10000 | 91.4 | 88.9 |
| SecBench | 89.7 | 72.5 |
| MMLU-Security | 89.7 | 78.2 |
Macro average — Saqr across all nine, Foundation-Sec-8B-Reasoning across the eight it reports
CyberMetric-2000
accuracyCybersecurity knowledge, 2,000 questions
96.2+11.9Foundation-Sec-8B-Reasoning8B84.3Cisco's own security-tuned model, at well under a third of the parameters.
CTIBench-ATE
accuracyMITRE ATT&CK technique extraction · no-think mode
72.1+23.0Foundation-Sec-8B-Reasoning8B49.1The largest margin of any benchmark here — a 60-sample set, reported with that size in mind.
MMLU-Security
accuracyComputer-security subset of MMLU
89.7+11.5Foundation-Sec-8B-Reasoning8B78.2Ahead of Cisco's own published figure on general security knowledge.
CTIBench-RCM
accuracyCVE-to-CWE root-cause mapping
77.8+9.0MVMinerva (Llama-3.1-8B)8B68.8Ahead of the RL-tuned Minerva model's published figure on the same task.
SecEval
accuracyMultiple-choice cybersecurity knowledge, nine domains
94.4+2.1GPT-5undisclosed92.3Ahead of OpenAI's frontier flagship, not just a cybersecurity specialist.
CyberMetric-2000
accuracyCybersecurity knowledge, 2,000 questions
96.2+3.6GPT-OSS-120B120B92.6An open-weight frontier model at more than four times Saqr's own parameter count.
Benchmark by benchmark
Benchmark by benchmark, against the security field
One panel per benchmark. Each shows every cybersecurity-specialised model that reports it — Foundation-Sec-8B-Reasoning, its non-reasoning sibling Foundation-Sec-8B-Instruct, and the research model Minerva — plus the general-purpose models Saqr is ahead of on that benchmark.
- #1 of 3
CTIBench-MCQA
Multiple-choice CTI knowledge across threat frameworks and taxonomies
SaqrR78.5- Foundation-Sec-8B-ReasoningR69.1
- Foundation-Sec-8B-Instruct65.0
General models Saqr leads here
- GPT-4.176.0
- GPT-5-MiniR75.3
- o3-miniR71.6
- GPT-OSS-120BR71.4
- ChatGPT-471.0
- Llama-3.3-70B-Instruct69.2
- GPT-5-NanoR68.8
- MSPhi-465.8
- Llama 3-70B65.7
- GPT-OSS-20BR65.5
- Gemini-1.565.4
- Llama 3-8B61.3
- Llama-3.1-8B-Instruct60.7
- ChatGPT-3.554.1
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 4
CTIBench-RCM
Maps CVE vulnerability descriptions to their root-cause CWE entry
SaqrR77.8- Foundation-Sec-8B-ReasoningR75.3
- Foundation-Sec-8B-Instruct70.4
- MVMinerva (Llama-3.1-8B)68.8
General models Saqr leads here
- GPT-4.173.0
- GPT-5R72.8
- GPT-5-MiniR72.3
- ChatGPT-472.0
- GPT-OSS-120BR71.2
- o3-miniR70.8
- Llama-3.3-70B-Instruct68.4
- GPT-5-NanoR67.2
- ChatGPT-3.567.2
- Gemini-1.566.6
- Llama 3-70B65.9
- MSPhi-462.9
- GPT-OSS-20BR61.0
- Llama-3.1-8B-Instruct53.1
- Llama 3-8B44.7
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 4
CTIBench-VSP
Predicts CVSS vulnerability severity from a text description
SaqrR90.1- MVMinerva (Llama-3.1-8B)87.6
- Foundation-Sec-8B-ReasoningR85.6
- Foundation-Sec-8B-Instruct84.0
General models Saqr leads here
- GPT-5-MiniR89.2
- GPT-OSS-120BR88.3
- GPT-OSS-20BR86.4
- GPT-4.184.8
- o3-miniR84.3
- Llama-3.3-70B-Instruct84.1
- GPT-5-NanoR82.2
- Llama-3.1-8B-Instruct81.1
- MSPhi-464.7
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 4
CTIBench-ATE
Extracts MITRE ATT&CK techniques from threat reports · no-think mode
SaqrR72.1- Foundation-Sec-8B-ReasoningR49.1
- MVMinerva (Llama-3.1-8B)48.4
- Foundation-Sec-8B-Instruct35.8
General models Saqr leads here
- GPT-4.169.6
- GPT-5-MiniR68.1
- o3-miniR59.9
- GPT-5R57.8
- Llama-3.3-70B-Instruct51.9
- GPT-OSS-20BR47.8
- GPT-5-NanoR45.3
- MSPhi-443.5
- GPT-OSS-120BR28.2
- Llama-3.1-8B-Instruct13.2
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 3
SecEval
Multiple-choice cybersecurity knowledge across nine security domains
SaqrR94.4- Foundation-Sec-8B-ReasoningR84.8
- Foundation-Sec-8B-Instruct82.9
General models Saqr leads here
- GPT-5R92.3
- GPT-4.191.9
- GPT-5-MiniR91.1
- o3-miniR90.8
- Llama-3.3-70B-Instruct90.6
- GPT-OSS-120BR90.4
- MSPhi-489.8
- GPT-5-NanoR88.4
- GPT-OSS-20BR87.0
- Llama-3.1-8B-Instruct83.2
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 4
CyberMetric-2000
Cybersecurity knowledge questions sourced from standards and RFCs
SaqrR96.2- Foundation-Sec-8B-Instruct84.7
- Foundation-Sec-8B-ReasoningR84.3
- MVMinerva (Llama-3.1-8B)84.2
General models Saqr leads here
- GPT-5R94.1
- GPT-4.193.7
- GPT-5-MiniR93.2
- o3-miniR93.0
- Llama-3.3-70B-Instruct93.0
- GPT-OSS-120BR92.6
- GPT-5-NanoR91.8
- MSPhi-491.2
- GPT-OSS-20BR89.3
- Llama-3.1-8B-Instruct85.1
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 1
CyberMetric-10000
The full 10,000-question CyberMetric knowledge set
SaqrR91.4
General models Saqr leads here
- GPT-488.9
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 3
SecBench
Multi-dimensional cybersecurity knowledge and reasoning questions
SaqrR89.7- Foundation-Sec-8B-Instruct74.4
- Foundation-Sec-8B-ReasoningR72.5
General models Saqr leads here
- YIYi-1.5-34B89.6
- GPT-5R88.8
- GPT-5-MiniR87.7
- GPT-4.187.2
- o3-miniR86.9
- Mixtral-8x7B86.1
- GPT-OSS-120BR85.3
- GLGLM-4-9B84.6
- Llama-3.3-70B-Instruct84.2
- GPT-5-NanoR83.8
- MSPhi-481.3
- GPT-OSS-20BR80.4
- DeepSeek-V2-Lite79.1
- Llama-3.1-8B-Instruct74.9
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
- #1 of 3
MMLU-Security
Computer-security subset of the MMLU knowledge benchmark
SaqrR89.7- Foundation-Sec-8B-ReasoningR78.2
- Foundation-Sec-8B-Instruct77.0
General models Saqr leads here
- GPT-5-MiniR88.4
- GPT-OSS-120BR88.0
- GPT-4.187.2
- GPT-OSS-20BR87.0
- Llama-3.3-70B-Instruct86.4
- o3-miniR85.8
- MSPhi-484.4
- GPT-5-NanoR83.6
- Llama-3.1-8B-Instruct76.8
accuracy · Saqr measured, others published. Specialists in full; general models shown where Saqr is ahead — the complete field is in the table above.
Panels are derived from the same data as the table below. The specialist field is shown in full; general models appear where Saqr is ahead of them, and the complete field is in the table further down. Saqr is measured; every other figure is published. CTIBench-RCM-2021 is tracked internally as a generalisation check — it carries no bar and no competitor figure, so it is not one of the nine panels here.
The frontier
CyberMetric-2000 accuracy, by parameter count
Plotting CyberMetric-2000 accuracy against parameter count shows what a table can only imply: Saqr leads every model plotted here, general frontier and cybersecurity specialist alike — including several with two to four times its own parameter count.
9 models plot here. A point needs both a published parameter count and a reported CyberMetric-2000 score — every other model in the table below is missing one or the other for this specific benchmark, so this stays an honest subset rather than a padded one.
Saqr leads accuracy at every parameter scale shown here, including against GPT-OSS-120B and Llama-3.3-70B-Instruct — both several times its own size. This chart is not an efficiency claim in Saqr’s favour; it plots a real relationship, in whichever direction the numbers fall.
Full competitive context
Every model with a real published figure on any of Saqr's nine gated benchmarks, sourced from each benchmark's own paper or the model's own technical report — OpenAI’s GPT-4 through GPT-5 and o3-mini, Google’s Gemini, Meta’s Llama 3, Microsoft’s Phi-4, Mistral’s Mixtral, DeepSeek and Zhipu’s GLM among them. Coverage is uneven by design — most of these were only tested on a subset of the nine — and a cell is left blank rather than filled with an estimate wherever a model was not tested on that particular benchmark.
Only Saqr’s row is measured by us; every other figure is that vendor’s or a benchmark paper’s published number, under protocols that differ and are mostly unstated. General-purpose frontier models (GPT-5, Gemini, Claude, Llama 4, DeepSeek) do not appear because none has a verified published score on any of these nine benchmarks — GPT-4’s single CyberMetric-10000 figure is the one exception.
Why it leads
What the numbers rest on
Ahead of every named specialist on 9 of 9 gated benchmarks
Foundation-Sec-8B-Reasoning, its non-reasoning sibling Foundation-Sec-8B-Instruct, and the research model Minerva (Llama-3.1-8B) are the cybersecurity-specialist models with published figures in this comparison. Saqr's published margin holds on every one of the nine benchmarks any of them reports — the full breakdown is benchmark by benchmark below.
Ahead of general frontier models, not just cybersecurity specialists
On CyberMetric-2000 and SecEval, Saqr's published margin holds against every model in this comparison — including OpenAI's GPT-5 and GPT-4.1, and GPT-OSS-120B, an open-weight model at more than four times Saqr's own parameter count. This isn't a cybersecurity-niche win; it's a real result against the frontier.
Beats a specialist's own repeated-sampling average
Foundation-Sec-8B-Reasoning's own technical report averages CyberMetric-2000, SecBench, SecEval and MMLU-Security over five sampled trials. Saqr's figures on those same four benchmarks are still ahead of that averaged number, not a single favourable run.
One 27.8B model, both reasoning modes
Saqr answers in thinking or non-thinking mode from the same dense 64-layer weights — fast responses for routine triage, or extended reasoning for complex vulnerability analysis, from one model.
Full-coverage evaluation, no subsampling
Every figure on this page is scored across a benchmark's complete question set — CyberMetric-10000 is the full ten thousand, not a sampled slice — and re-run after every training iteration until it clears its bar. 9 of 9 currently do.
The flagship
Saqr by ZySec
Infinia Technologies' flagship cybersecurity model — built by ZySec, Infinia Technologies' cybersecurity arm.
Built for security teams
Full-coverage evaluation, no subsampling
Every benchmark on this page is scored across its complete question set — CyberMetric-10000 is all ten thousand questions.
Purpose-built for cybersecurity
Trained specifically for cyber threat intelligence, vulnerability analysis and security knowledge — not adapted from a general-purpose assistant.
Deploy on your own infrastructure
Full model weights, no API dependency — run it air-gapped or in your own cloud, under your own controls.
Access
Evaluate it, or tell us what you need
Tell us about your organisation and what you need — we'll follow up about deployment options, evaluation access, and pricing.
- Full model weights in bfloat16, text and image input — deployable entirely on your own infrastructure.
- The complete model card, with full benchmark methodology and results.
- Our evaluation setup, so a reviewer can reproduce every figure here.