Fikra Benchmark Phase 1
Overview
Fikra Benchmark Phase 1 reports an author-conducted evaluation of four language-model configurations served through the Fikra API.
The study evaluates the configurations using a frozen 500-task suite assembled from MMLU-Pro, GPQA-Diamond, GSM8K, IFEval, LiveCodeBench, BFCL, and a Fikra-owned FRES set.
The evaluation was designed around a simple methodological principle: capability measurements and system-execution measurements should be treated as related but distinct observations.
Different benchmark families examine different properties. Scientific reasoning, mathematical problem solving, instruction following, and structured tool invocation do not measure the same capability. A deployed API also introduces operational characteristics such as latency, failed executions, response truncation, and evaluator compatibility.
For that reason, Phase 1 does not attempt to collapse the experiment into one universal score.
The experiment generated 2,000 planned model-task evaluations. Of those, 1,894 completed successfully, representing 94.7% successful execution coverage. A total of 395 of the 500 frozen tasks had successful results for all four configurations. These coverage figures include tasks from LiveCodeBench and FRES, even though those two components were later excluded from the final scored analysis.
The completed scored analysis covers GPQA-Diamond, GSM8K, MMLU-Pro, IFEval, and BFCL. LiveCodeBench and FRES are not scored because reliable evaluation environments were not established during the frozen experiment.
The author operates the Fikra API through which all evaluated configurations were served, and the FRES evaluation set is Fikra-owned. This disclosure is material to interpretation of the results.
Research Questions
The study addresses four primary questions.
-
How do the four configurations differ across heterogeneous capability categories?
-
How complete and comparable are the resulting measurements when failed model-task executions are treated separately from incorrect answers?
-
What execution characteristics accompany the observed capability measurements?
-
What methodological limitations arise when public benchmark evaluators are reproduced in an API-based experimental environment?
Contributions
Phase 1 contributes:
- a documented frozen 500-task evaluation suite
- a reproducible four-configuration evaluation record
- capability measurements across several benchmark families
- execution measurements covering successful-call coverage and latency
- an explicit accounting of incomplete evaluations
- documentation of evaluator compatibility issues
- research artifacts containing the benchmark manifest, raw records, diagnostics, and analysis material
The experiment also provides an empirical basis for later work on task-aware model selection and routing. A routing system itself is not evaluated in Phase 1.
Evaluation Landscape
The benchmark intentionally combines evaluation families with different semantics.
MMLU-Pro
MMLU-Pro is used as a broad, reasoning-focused multiple-choice evaluation. Its questions use ten answer choices, A through J.
GPQA-Diamond
GPQA targets difficult graduate-level scientific questions across biology, physics, and chemistry.
GSM8K
GSM8K evaluates multi-step mathematical reasoning using grade-school mathematics problems.
IFEval
IFEval evaluates verifiable instruction-following properties such as required phrases, formatting constraints, and length constraints. Its design emphasizes deterministic automatic verification rather than subjective judging.
BFCL
BFCL evaluates structured function calling. Phase 1 uses AST-based checking and includes serial and parallel function-calling categories.
LiveCodeBench
LiveCodeBench was included in the frozen benchmark design as the coding component. It was not included in the final scored analysis because a reliable execution environment could not be established.
FRES
FRES is a Fikra-owned evaluation set. It was retained in the frozen benchmark definition but excluded from final scoring because reliable evaluation conditions were not established during Phase 1.
The deliberately heterogeneous benchmark design means that a result from one evaluation family is not treated as a proxy for another.
Benchmark Design
The benchmark was frozen on 26 September 2026.
The frozen suite contained 500 tasks.
| Dataset | Tasks |
|---|---|
| MMLU-Pro | 100 |
| GPQA-Diamond | 50 |
| GSM8K | 75 |
| IFEval | 75 |
| LiveCodeBench | 100 |
| BFCL | 50 |
| FRES | 50 |
| Total | 500 |
The resulting benchmark SHA-256 is:
bf904af6e4a7f909c92642f2299f1001c7f1aff88894684fd61f2e799df72bfb
The experiment identifier is:
fikra-phase1-2026-09-26-v1
The run is identified as:
fikra-benchmark-phase1-2026-09-26-4model
The evaluation unit is one model configuration paired with one frozen task. Four configurations applied to 500 tasks yield 2,000 planned model-task evaluations.
Restricted source material, including gated benchmark question text where applicable, is not reproduced in the public manuscript or research page.
Evaluated Configurations
The evaluated Fikra configurations were:
| Fikra configuration | Upstream model |
|---|---|
| Fikra Flash | GPT-OSS 20B |
| Fikra Pro 20B | Qwen 3.8 27B |
| Fikra Pro 120B | GPT-OSS 120B |
| Fikra Qwen Max | Qwen 3.8 Max |
These configuration names are Fikra product labels. They should not be interpreted as upstream parameter-count labels.
The two GPT-OSS-based configurations use the open-weight models identified in the paper's references.
The configurations are evaluated as exposed by the Fikra API, including serving-layer processing. Consequently, the measurements describe the complete API configuration under test rather than a provider-independent intrinsic measurement of the upstream model. :chatgpt-content-reference{index="2"}
Generation Protocol
Unless an evaluator required otherwise, requests used:
- temperature = 0
- top-p = 1
- one attempt per model-task pair
- maximum output length = 4,096 tokens
- identical frozen task wording across configurations
- no external retrieval
- no human intervention
The raw experiment record stores task identifiers, dataset information, model identifiers, response data, timing information, token usage where available, HTTP/API status, and failure information. :chatgpt-content-reference{index="3"}
Evaluation Coverage
A successful generation and a correct answer are separate concepts.
A failed or missing model-task execution is therefore not counted as an incorrect answer.
Phase 1 completed 1,894 of 2,000 planned evaluations, producing 94.7% successful execution coverage.
| Configuration | Successful | Planned | Coverage |
|---|---|---|---|
| Fikra Flash | 499 | 500 | 99.8% |
| Fikra Pro 20B | 395 | 500 | 79.0% |
| Fikra Pro 120B | 500 | 500 | 100.0% |
| Fikra Qwen Max | 500 | 500 | 100.0% |
| Total | 1,894 | 2,000 | 94.7% |
There were 395 frozen tasks for which all four configurations produced successful results.
These 395 tasks span the complete 500-task frozen design, including the 150 LiveCodeBench and FRES tasks that were ultimately excluded from scored analysis. They therefore represent generation coverage rather than coverage of the scored benchmark components.
The incomplete coverage is concentrated in Fikra Pro 20B and one Fikra Flash evaluation.
MMLU-Pro requires particular caution. The Pro 20B configuration produced only 31 successful evaluations for that dataset, of which 21 were correct. The paper therefore reports the result descriptively rather than treating it as equivalent to a complete 100-task evaluation. :chatgpt-content-reference{index="4"}
Scoring Methodology
Phase 1 uses deterministic or benchmark-native scoring wherever available.
GPQA-Diamond
Exact answer-choice matching.
GSM8K
Normalized final-answer matching.
MMLU-Pro
Answer-choice matching across all ten answer options, A through J.
IFEval
The benchmark's IFEval verification logic was retained.
During evaluation, 116 instruction entries contained keyword arguments that were not accepted by individual evaluator signatures. Compatibility-oriented keyword-argument sanitization was therefore applied.
The underlying verification logic was retained, but the resulting installation should not be characterized as an untouched execution of the upstream evaluator.
BFCL
BFCL scoring uses the AST-checking logic of the BFCL codebase, pinned to commit:
f7cf735
The complete BFCL evaluator stack encountered dependency and API compatibility problems. Final BFCL scoring therefore used the AST-checking logic in standalone form rather than the complete provider-handler and tree-sitter stack. :chatgpt-content-reference{index="5"}
Excluded Components
LiveCodeBench and FRES were intentionally excluded from the final scored analysis.
No score was estimated, imputed, or inferred for either component.
The frozen benchmark definition retains both components as part of the original 500-task experiment, but the final capability analysis excludes them because reliable evaluation environments could not be established. :chatgpt-content-reference{index="6"}
Results
The final scored analysis covers five components:
- GPQA-Diamond
- GSM8K
- MMLU-Pro
- IFEval
- BFCL
The results are reported by dataset rather than collapsed into a single aggregate score.
This is intentional. The benchmark components have different semantics, and evaluation coverage is incomplete.
Capability Results
| Dataset | Fikra Flash | Fikra Pro 20B | Fikra Pro 120B | Fikra Qwen Max |
|---|---|---|---|---|
| GPQA-Diamond | 48.00% | 46.00% | 66.00% | 88.00% |
| GSM8K | 85.33% | 84.00% | 84.00% | 94.67% |
| MMLU-Pro | 77.00% | 67.74%* | 82.00% | 90.00% |
| IFEval (strict) | 68.00% | 76.00% | 72.00% | 86.67% |
| IFEval (loose) | 69.33% | 78.67% | 76.00% | 88.00% |
| BFCL | 54.00% | 86.00% | 58.00% | 82.00% |
* The Pro 20B MMLU-Pro result is based on 21 correct responses from 31 successful evaluations rather than a complete 100-task run.
The observed score ranges are 46–88% for GPQA-Diamond, 84–94.67% for GSM8K, 67.74–90% for MMLU-Pro, 68–86.67% for strict IFEval, and 54–86% for BFCL. :chatgpt-content-reference{index="7"}
The results show heterogeneous capability profiles across the four configurations. A configuration can have materially different behavior across reasoning, mathematics, instruction following, and function calling.
The study therefore does not treat the completed results as evidence for one universal ordering.
The paper also notes that sample sizes are relatively small for several components. At sample sizes of 50, 75, and 100, approximate 95% Wilson intervals around observed accuracies can extend by roughly ±13, ±11, and ±10 percentage points respectively. Comparisons involving the incompletely covered Pro 20B configuration require particular caution because it is evaluated on a different and potentially easier subset of tasks. :chatgpt-content-reference{index="8"}
Function Calling
BFCL reveals substantial category variation beneath the aggregate function-calling scores.
The four evaluated categories were:
- simple Python
- multiple calls
- parallel calls
- parallel multiple calls
All four configurations achieved at least 90% on simple Python except Fikra Flash at 85%.
All four configurations achieved 100% on the multiple-call category.
The largest differences occurred in the parallel categories:
| Configuration | Parallel | Parallel Multiple |
|---|---|---|
| Fikra Flash | 0/10 | 0/10 |
| Fikra Pro 20B | 6/10 | 9/10 |
| Fikra Pro 120B | 0/10 | 1/10 |
| Fikra Qwen Max | 7/10 | 6/10 |
This provides a concrete reason to report structured tool-use behavior separately from conventional reasoning and mathematics benchmarks. :chatgpt-content-reference{index="9"}
Instruction Following
Strict IFEval scores were:
| Configuration | Strict | Loose |
|---|---|---|
| Fikra Flash | 68.00% | 69.33% |
| Fikra Pro 20B | 76.00% | 78.67% |
| Fikra Pro 120B | 72.00% | 76.00% |
| Fikra Qwen Max | 86.67% | 88.00% |
Strict and loose measurements are retained because they represent different outputs from the benchmark's verification procedure. :chatgpt-content-reference{index="10"}
Operational Performance
Phase 1 also measures execution characteristics.
These values describe the observed Fikra API conditions of the experiment rather than provider-independent intrinsic model latency.
Requests were issued concurrently using 12 workers at approximately 80 requests per minute. As a result, observed latency and time to first token can include queueing, retry, and rate-limit effects in addition to model inference time.
Median measurements over successful calls were:
| Configuration | Successful calls | TTFT observations | Median TTFT | Median latency | Median output tokens |
|---|---|---|---|---|---|
| Fikra Flash | 499 | 453 | 2,095.3 ms | 3,734.3 ms | 401 |
| Fikra Pro 20B | 395 | 385 | 1,050.8 ms | 2,703.9 ms | 240 |
| Fikra Pro 120B | 500 | 492 | 2,242.7 ms | 3,711.2 ms | 421 |
| Fikra Qwen Max | 500 | 500 | 13,289.5 ms | 16,065.4 ms | 529.5 |
Observed median total latency therefore ranged from 2.70 seconds to 16.07 seconds, while median TTFT ranged from 1.05 seconds to 13.29 seconds. :chatgpt-content-reference{index="11"} :chatgpt-content-reference{index="12"}
Methodological Findings
One of the clearest findings from the experiment concerns the distinction between reproducing a benchmark definition and reproducing its complete evaluator environment.
The frozen task set can be preserved and cryptographically hashed.
The evaluator can still depend on:
- package versions
- provider interfaces
- runtime assumptions
- generated response schemas
- execution dependencies
The IFEval compatibility issue demonstrates that an ostensibly deterministic public evaluator can still require interface-level adaptation.
The BFCL experience demonstrates a stronger form of environment dependence. The complete evaluator stack could not be executed in the available environment, while its AST-checking logic could still be applied independently to appropriate outputs.
The experiment therefore treats evaluator compatibility as part of the research record rather than hiding it.
Why Coverage Matters
The Phase 1 results also illustrate why evaluation coverage should be published alongside accuracy.
A missing response is an execution observation.
It is not evidence that the model produced an incorrect answer.
Conflating the two would mix infrastructure reliability with model capability.
That distinction is particularly important in API-based evaluations, where model serving and evaluation environments introduce failure modes that are not necessarily equivalent to model errors. :chatgpt-content-reference{index="13"}
Discussion
The completed results support three bounded observations.
First, the four configurations exhibit differentiated capability profiles across the measured task families.
Second, tool-use performance can differ substantially from conventional reasoning and mathematics performance, meaning that a single capability dimension is insufficient for deployment-oriented evaluation.
Third, execution characteristics such as latency and successful-call coverage are relevant empirical dimensions alongside benchmark accuracy.
The study does not establish that one configuration is universally superior.
It also does not establish that model size, upstream model family, or any other individual model characteristic causally explains the observed differences.
The experiment measures Fikra configurations as exposed through the Fikra API under the documented protocol. Provider-side implementation, serving infrastructure, prompt handling, and runtime conditions can all contribute to the observations.
The heterogeneous profiles motivate later research into task-aware routing.
A routing system could potentially use task characteristics and operational constraints when selecting among configurations, but that hypothesis requires separate experimental evaluation. Phase 1 does not evaluate a routing system and makes no claim about the quality, cost, or latency of such a system. :chatgpt-content-reference{index="14"}
Limitations
Incomplete Coverage
The experiment completed 1,894 of 2,000 planned evaluations, leaving 106 unsuccessful evaluations.
Fikra Pro 20B has materially lower coverage than the other configurations, and its MMLU-Pro result is based on only 31 successful evaluations.
Evaluator Compatibility
IFEval required compatibility-oriented keyword-argument sanitization.
BFCL required standalone use of its AST-checking logic rather than execution of the complete evaluator stack.
These decisions are documented precisely because they affect interpretation and reproducibility.
Excluded Benchmarks
LiveCodeBench and FRES were not scored.
No numerical value is inferred for either component.
Harness Validity
The raw records show that 126 of the 1,894 successful generations ended with the finish reason length.
The distribution was:
| Configuration | Length-terminated successful generations |
|---|---|
| Fikra Flash | 59 |
| Fikra Pro 20B | 37 |
| Fikra Pro 120B | 29 |
| Fikra Qwen Max | 1 |
| Total | 126 |
Those generations reached the 4,096-token output cap and may therefore lack a final answer.
The paper notes that this includes 17 of the 50 Fikra Pro 20B GPQA-Diamond generations.
Answer extraction also relies on simple letter-based rules that can mishandle verbose or truncated responses.
The near-zero results of the GPT-OSS-based configurations on the parallel BFCL categories may also reflect output-format handling rather than underlying capability.
Reasoning-effort behavior was not varied in Phase 1, and temperature-0 behavior was not compared against alternative settings.
The paper therefore treats the results as measurements of these configurations under this particular harness rather than estimates of upstream models' peak capability.
Manual verification of a sample of responses is identified as a priority for the next phase. :chatgpt-content-reference{index="15"}
No Aggregate Score
Phase 1 deliberately does not define an authoritative aggregate score.
Constructing one would require decisions about how heterogeneous capabilities should be weighted and how incomplete coverage should be handled.
The publication therefore reports dataset-native measurements and common-coverage facts rather than manufacturing a universal ranking.
Reproducibility and Research Artifacts
The research package contains:
- the frozen benchmark artifact
- the experiment manifest
- raw result files
- the production notebook
- dataset revision metadata
- evaluator diagnostics
- scoring records
- research claims
- reproducibility notes
The benchmark hash and experiment identifiers are recorded above.
Restricted benchmark content and credentials are not included in the public manuscript source.
The manuscript is intentionally smaller than the complete research archive because publicly posted manuscript source is itself downloadable.
The code, benchmark metadata, and raw results are identified in the paper as being available through the Fikra Benchmark repository, with the research also archived on Zenodo. :chatgpt-content-reference{index="16"}
Disclosure
This is author-conducted research.
James Miano operates the Fikra API through which all evaluated configurations were served, and the FRES set is Fikra-owned.
That relationship should be considered when interpreting the measurements.
The frozen benchmark hash and raw records are provided to support independent verification. :chatgpt-content-reference{index="17"}
Conclusion
Fikra Benchmark Phase 1 provides a documented comparative evaluation of four Fikra-served LLM configurations across heterogeneous capability and execution dimensions.
The frozen suite contained 500 tasks and produced 1,894 successful model-task evaluations.
The completed measurements show differentiated profiles across scientific reasoning, mathematics, instruction following, and function calling, while operational measurements show substantial differences in latency and execution coverage.
The central methodological conclusion is that capability measurements should be interpreted together with evaluation coverage, task semantics, and execution conditions.
Phase 1 therefore serves as an empirical foundation for subsequent work on task-aware model selection and routing rather than as a universal model leaderboard.
Publication Record
Publication: Fikra Benchmark Phase 1
Author: James Miano
Organization: Lacesse Ventures (operator of the Fikra API)
Publication date: September 2026
Archival record: Zenodo
DOI provided for the Fikra publication page: 10.5281/zenodo.23078946