Fikra Benchmark Phase 1

Overview

Fikra Benchmark Phase 1 reports an author-conducted evaluation of four language-model configurations served through the Fikra API.

The study evaluates the configurations using a frozen 500-task suite assembled from MMLU-Pro, GPQA-Diamond, GSM8K, IFEval, LiveCodeBench, BFCL, and a Fikra-owned FRES set.

The evaluation was designed around a simple methodological principle: capability measurements and system-execution measurements should be treated as related but distinct observations.

Different benchmark families examine different properties. Scientific reasoning, mathematical problem solving, instruction following, and structured tool invocation do not measure the same capability. A deployed API also introduces operational characteristics such as latency, failed executions, response truncation, and evaluator compatibility.

For that reason, Phase 1 does not attempt to collapse the experiment into one universal score.

The experiment generated 2,000 planned model-task evaluations. Of those, 1,894 completed successfully, representing 94.7% successful execution coverage. A total of 395 of the 500 frozen tasks had successful results for all four configurations. These coverage figures include tasks from LiveCodeBench and FRES, even though those two components were later excluded from the final scored analysis.

The completed scored analysis covers GPQA-Diamond, GSM8K, MMLU-Pro, IFEval, and BFCL. LiveCodeBench and FRES are not scored because reliable evaluation environments were not established during the frozen experiment.

The author operates the Fikra API through which all evaluated configurations were served, and the FRES evaluation set is Fikra-owned. This disclosure is material to interpretation of the results.

Research Questions

The study addresses four primary questions.

  1. How do the four configurations differ across heterogeneous capability categories?

  2. How complete and comparable are the resulting measurements when failed model-task executions are treated separately from incorrect answers?

  3. What execution characteristics accompany the observed capability measurements?

  4. What methodological limitations arise when public benchmark evaluators are reproduced in an API-based experimental environment?

Contributions

Phase 1 contributes:

  • a documented frozen 500-task evaluation suite
  • a reproducible four-configuration evaluation record
  • capability measurements across several benchmark families
  • execution measurements covering successful-call coverage and latency
  • an explicit accounting of incomplete evaluations
  • documentation of evaluator compatibility issues
  • research artifacts containing the benchmark manifest, raw records, diagnostics, and analysis material

The experiment also provides an empirical basis for later work on task-aware model selection and routing. A routing system itself is not evaluated in Phase 1.

Evaluation Landscape

The benchmark intentionally combines evaluation families with different semantics.

MMLU-Pro

MMLU-Pro is used as a broad, reasoning-focused multiple-choice evaluation. Its questions use ten answer choices, A through J.

GPQA-Diamond

GPQA targets difficult graduate-level scientific questions across biology, physics, and chemistry.

GSM8K

GSM8K evaluates multi-step mathematical reasoning using grade-school mathematics problems.

IFEval

IFEval evaluates verifiable instruction-following properties such as required phrases, formatting constraints, and length constraints. Its design emphasizes deterministic automatic verification rather than subjective judging.

BFCL

BFCL evaluates structured function calling. Phase 1 uses AST-based checking and includes serial and parallel function-calling categories.

LiveCodeBench

LiveCodeBench was included in the frozen benchmark design as the coding component. It was not included in the final scored analysis because a reliable execution environment could not be established.

FRES

FRES is a Fikra-owned evaluation set. It was retained in the frozen benchmark definition but excluded from final scoring because reliable evaluation conditions were not established during Phase 1.

The deliberately heterogeneous benchmark design means that a result from one evaluation family is not treated as a proxy for another.

Benchmark Design

The benchmark was frozen on 26 September 2026.

The frozen suite contained 500 tasks.

Dataset Tasks
MMLU-Pro 100
GPQA-Diamond 50
GSM8K 75
IFEval 75
LiveCodeBench 100
BFCL 50
FRES 50
Total 500

The resulting benchmark SHA-256 is:

bf904af6e4a7f909c92642f2299f1001c7f1aff88894684fd61f2e799df72bfb

The experiment identifier is:

fikra-phase1-2026-09-26-v1

The run is identified as:

fikra-benchmark-phase1-2026-09-26-4model

The evaluation unit is one model configuration paired with one frozen task. Four configurations applied to 500 tasks yield 2,000 planned model-task evaluations.

Restricted source material, including gated benchmark question text where applicable, is not reproduced in the public manuscript or research page.

Evaluated Configurations

The evaluated Fikra configurations were:

Fikra configuration Upstream model
Fikra Flash GPT-OSS 20B
Fikra Pro 20B Qwen 3.8 27B
Fikra Pro 120B GPT-OSS 120B
Fikra Qwen Max Qwen 3.8 Max

These configuration names are Fikra product labels. They should not be interpreted as upstream parameter-count labels.

The two GPT-OSS-based configurations use the open-weight models identified in the paper's references.

The configurations are evaluated as exposed by the Fikra API, including serving-layer processing. Consequently, the measurements describe the complete API configuration under test rather than a provider-independent intrinsic measurement of the upstream model. :chatgpt-content-reference{index="2"}

Generation Protocol

Unless an evaluator required otherwise, requests used:

  • temperature = 0
  • top-p = 1
  • one attempt per model-task pair
  • maximum output length = 4,096 tokens
  • identical frozen task wording across configurations
  • no external retrieval
  • no human intervention

The raw experiment record stores task identifiers, dataset information, model identifiers, response data, timing information, token usage where available, HTTP/API status, and failure information. :chatgpt-content-reference{index="3"}

Evaluation Coverage

A successful generation and a correct answer are separate concepts.

A failed or missing model-task execution is therefore not counted as an incorrect answer.

Phase 1 completed 1,894 of 2,000 planned evaluations, producing 94.7% successful execution coverage.

Configuration Successful Planned Coverage
Fikra Flash 499 500 99.8%
Fikra Pro 20B 395 500 79.0%
Fikra Pro 120B 500 500 100.0%
Fikra Qwen Max 500 500 100.0%
Total 1,894 2,000 94.7%

There were 395 frozen tasks for which all four configurations produced successful results.

These 395 tasks span the complete 500-task frozen design, including the 150 LiveCodeBench and FRES tasks that were ultimately excluded from scored analysis. They therefore represent generation coverage rather than coverage of the scored benchmark components.

The incomplete coverage is concentrated in Fikra Pro 20B and one Fikra Flash evaluation.

MMLU-Pro requires particular caution. The Pro 20B configuration produced only 31 successful evaluations for that dataset, of which 21 were correct. The paper therefore reports the result descriptively rather than treating it as equivalent to a complete 100-task evaluation. :chatgpt-content-reference{index="4"}

Scoring Methodology

Phase 1 uses deterministic or benchmark-native scoring wherever available.

GPQA-Diamond

Exact answer-choice matching.

GSM8K

Normalized final-answer matching.

MMLU-Pro

Answer-choice matching across all ten answer options, A through J.

IFEval

The benchmark's IFEval verification logic was retained.

During evaluation, 116 instruction entries contained keyword arguments that were not accepted by individual evaluator signatures. Compatibility-oriented keyword-argument sanitization was therefore applied.

The underlying verification logic was retained, but the resulting installation should not be characterized as an untouched execution of the upstream evaluator.

BFCL

BFCL scoring uses the AST-checking logic of the BFCL codebase, pinned to commit:

f7cf735

The complete BFCL evaluator stack encountered dependency and API compatibility problems. Final BFCL scoring therefore used the AST-checking logic in standalone form rather than the complete provider-handler and tree-sitter stack. :chatgpt-content-reference{index="5"}

Excluded Components

LiveCodeBench and FRES were intentionally excluded from the final scored analysis.

No score was estimated, imputed, or inferred for either component.

The frozen benchmark definition retains both components as part of the original 500-task experiment, but the final capability analysis excludes them because reliable evaluation environments could not be established. :chatgpt-content-reference{index="6"}

Results

The final scored analysis covers five components:

  • GPQA-Diamond
  • GSM8K
  • MMLU-Pro
  • IFEval
  • BFCL

The results are reported by dataset rather than collapsed into a single aggregate score.

This is intentional. The benchmark components have different semantics, and evaluation coverage is incomplete.

Capability Results

Dataset Fikra Flash Fikra Pro 20B Fikra Pro 120B Fikra Qwen Max
GPQA-Diamond 48.00% 46.00% 66.00% 88.00%
GSM8K 85.33% 84.00% 84.00% 94.67%
MMLU-Pro 77.00% 67.74%* 82.00% 90.00%
IFEval (strict) 68.00% 76.00% 72.00% 86.67%
IFEval (loose) 69.33% 78.67% 76.00% 88.00%
BFCL 54.00% 86.00% 58.00% 82.00%

* The Pro 20B MMLU-Pro result is based on 21 correct responses from 31 successful evaluations rather than a complete 100-task run.

The observed score ranges are 46–88% for GPQA-Diamond, 84–94.67% for GSM8K, 67.74–90% for MMLU-Pro, 68–86.67% for strict IFEval, and 54–86% for BFCL. :chatgpt-content-reference{index="7"}

The results show heterogeneous capability profiles across the four configurations. A configuration can have materially different behavior across reasoning, mathematics, instruction following, and function calling.

The study therefore does not treat the completed results as evidence for one universal ordering.

The paper also notes that sample sizes are relatively small for several components. At sample sizes of 50, 75, and 100, approximate 95% Wilson intervals around observed accuracies can extend by roughly ±13, ±11, and ±10 percentage points respectively. Comparisons involving the incompletely covered Pro 20B configuration require particular caution because it is evaluated on a different and potentially easier subset of tasks. :chatgpt-content-reference{index="8"}

Function Calling

BFCL reveals substantial category variation beneath the aggregate function-calling scores.

The four evaluated categories were:

  • simple Python
  • multiple calls
  • parallel calls
  • parallel multiple calls

All four configurations achieved at least 90% on simple Python except Fikra Flash at 85%.

All four configurations achieved 100% on the multiple-call category.

The largest differences occurred in the parallel categories:

Configuration Parallel Parallel Multiple
Fikra Flash 0/10 0/10
Fikra Pro 20B 6/10 9/10
Fikra Pro 120B 0/10 1/10
Fikra Qwen Max 7/10 6/10

This provides a concrete reason to report structured tool-use behavior separately from conventional reasoning and mathematics benchmarks. :chatgpt-content-reference{index="9"}

Instruction Following

Strict IFEval scores were:

Configuration Strict Loose
Fikra Flash 68.00% 69.33%
Fikra Pro 20B 76.00% 78.67%
Fikra Pro 120B 72.00% 76.00%
Fikra Qwen Max 86.67% 88.00%

Strict and loose measurements are retained because they represent different outputs from the benchmark's verification procedure. :chatgpt-content-reference{index="10"}

Operational Performance

Phase 1 also measures execution characteristics.

These values describe the observed Fikra API conditions of the experiment rather than provider-independent intrinsic model latency.

Requests were issued concurrently using 12 workers at approximately 80 requests per minute. As a result, observed latency and time to first token can include queueing, retry, and rate-limit effects in addition to model inference time.

Median measurements over successful calls were:

Configuration Successful calls TTFT observations Median TTFT Median latency Median output tokens
Fikra Flash 499 453 2,095.3 ms 3,734.3 ms 401
Fikra Pro 20B 395 385 1,050.8 ms 2,703.9 ms 240
Fikra Pro 120B 500 492 2,242.7 ms 3,711.2 ms 421
Fikra Qwen Max 500 500 13,289.5 ms 16,065.4 ms 529.5

Observed median total latency therefore ranged from 2.70 seconds to 16.07 seconds, while median TTFT ranged from 1.05 seconds to 13.29 seconds. :chatgpt-content-reference{index="11"} :chatgpt-content-reference{index="12"}

Methodological Findings

One of the clearest findings from the experiment concerns the distinction between reproducing a benchmark definition and reproducing its complete evaluator environment.

The frozen task set can be preserved and cryptographically hashed.

The evaluator can still depend on:

  • package versions
  • provider interfaces
  • runtime assumptions
  • generated response schemas
  • execution dependencies

The IFEval compatibility issue demonstrates that an ostensibly deterministic public evaluator can still require interface-level adaptation.

The BFCL experience demonstrates a stronger form of environment dependence. The complete evaluator stack could not be executed in the available environment, while its AST-checking logic could still be applied independently to appropriate outputs.

The experiment therefore treats evaluator compatibility as part of the research record rather than hiding it.

Why Coverage Matters

The Phase 1 results also illustrate why evaluation coverage should be published alongside accuracy.

A missing response is an execution observation.

It is not evidence that the model produced an incorrect answer.

Conflating the two would mix infrastructure reliability with model capability.

That distinction is particularly important in API-based evaluations, where model serving and evaluation environments introduce failure modes that are not necessarily equivalent to model errors. :chatgpt-content-reference{index="13"}

Discussion

The completed results support three bounded observations.

First, the four configurations exhibit differentiated capability profiles across the measured task families.

Second, tool-use performance can differ substantially from conventional reasoning and mathematics performance, meaning that a single capability dimension is insufficient for deployment-oriented evaluation.

Third, execution characteristics such as latency and successful-call coverage are relevant empirical dimensions alongside benchmark accuracy.

The study does not establish that one configuration is universally superior.

It also does not establish that model size, upstream model family, or any other individual model characteristic causally explains the observed differences.

The experiment measures Fikra configurations as exposed through the Fikra API under the documented protocol. Provider-side implementation, serving infrastructure, prompt handling, and runtime conditions can all contribute to the observations.

The heterogeneous profiles motivate later research into task-aware routing.

A routing system could potentially use task characteristics and operational constraints when selecting among configurations, but that hypothesis requires separate experimental evaluation. Phase 1 does not evaluate a routing system and makes no claim about the quality, cost, or latency of such a system. :chatgpt-content-reference{index="14"}

Limitations

Incomplete Coverage

The experiment completed 1,894 of 2,000 planned evaluations, leaving 106 unsuccessful evaluations.

Fikra Pro 20B has materially lower coverage than the other configurations, and its MMLU-Pro result is based on only 31 successful evaluations.

Evaluator Compatibility

IFEval required compatibility-oriented keyword-argument sanitization.

BFCL required standalone use of its AST-checking logic rather than execution of the complete evaluator stack.

These decisions are documented precisely because they affect interpretation and reproducibility.

Excluded Benchmarks

LiveCodeBench and FRES were not scored.

No numerical value is inferred for either component.

Harness Validity

The raw records show that 126 of the 1,894 successful generations ended with the finish reason length.

The distribution was:

Configuration Length-terminated successful generations
Fikra Flash 59
Fikra Pro 20B 37
Fikra Pro 120B 29
Fikra Qwen Max 1
Total 126

Those generations reached the 4,096-token output cap and may therefore lack a final answer.

The paper notes that this includes 17 of the 50 Fikra Pro 20B GPQA-Diamond generations.

Answer extraction also relies on simple letter-based rules that can mishandle verbose or truncated responses.

The near-zero results of the GPT-OSS-based configurations on the parallel BFCL categories may also reflect output-format handling rather than underlying capability.

Reasoning-effort behavior was not varied in Phase 1, and temperature-0 behavior was not compared against alternative settings.

The paper therefore treats the results as measurements of these configurations under this particular harness rather than estimates of upstream models' peak capability.

Manual verification of a sample of responses is identified as a priority for the next phase. :chatgpt-content-reference{index="15"}

No Aggregate Score

Phase 1 deliberately does not define an authoritative aggregate score.

Constructing one would require decisions about how heterogeneous capabilities should be weighted and how incomplete coverage should be handled.

The publication therefore reports dataset-native measurements and common-coverage facts rather than manufacturing a universal ranking.

Reproducibility and Research Artifacts

The research package contains:

  • the frozen benchmark artifact
  • the experiment manifest
  • raw result files
  • the production notebook
  • dataset revision metadata
  • evaluator diagnostics
  • scoring records
  • research claims
  • reproducibility notes

The benchmark hash and experiment identifiers are recorded above.

Restricted benchmark content and credentials are not included in the public manuscript source.

The manuscript is intentionally smaller than the complete research archive because publicly posted manuscript source is itself downloadable.

The code, benchmark metadata, and raw results are identified in the paper as being available through the Fikra Benchmark repository, with the research also archived on Zenodo. :chatgpt-content-reference{index="16"}

Disclosure

This is author-conducted research.

James Miano operates the Fikra API through which all evaluated configurations were served, and the FRES set is Fikra-owned.

That relationship should be considered when interpreting the measurements.

The frozen benchmark hash and raw records are provided to support independent verification. :chatgpt-content-reference{index="17"}

Conclusion

Fikra Benchmark Phase 1 provides a documented comparative evaluation of four Fikra-served LLM configurations across heterogeneous capability and execution dimensions.

The frozen suite contained 500 tasks and produced 1,894 successful model-task evaluations.

The completed measurements show differentiated profiles across scientific reasoning, mathematics, instruction following, and function calling, while operational measurements show substantial differences in latency and execution coverage.

The central methodological conclusion is that capability measurements should be interpreted together with evaluation coverage, task semantics, and execution conditions.

Phase 1 therefore serves as an empirical foundation for subsequent work on task-aware model selection and routing rather than as a universal model leaderboard.

Publication Record

Publication: Fikra Benchmark Phase 1

Author: James Miano

Organization: Lacesse Ventures (operator of the Fikra API)

Publication date: September 2026

Archival record: Zenodo

DOI provided for the Fikra publication page: 10.5281/zenodo.23078946