Benchmark
Fikra Benchmark Phase 1: Comparative Evaluation of Four LLM Configurations Across Reasoning, Instruction Following, Mathematics, and Tool Use
An author-conducted evaluation of four language-model configurations served through the Fikra API, using a frozen 500-task suite to examine reasoning, mathematics, instruction following, function calling, execution coverage, latency, and evaluation methodology.
FI.