GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 2
Unlocking the Potential of Small-to-Medium LLMs in Autonomous Agent Architectures: A Comparative Analysis of Gemini, Llama, and Qwen
Authors
Gurleen Singh, Rizwan Yousuf
Abstract
The landscape of Artificial Intelligence is experiencing a paradigm shift from traditional conversational chatbots to robust, proactive autonomous agents capable of complex tool orchestration, multi-step reasoning, and dynamic decision-making. However, deploying state-of-the-art agentic systems is fundamentally bottlenecked by reliance on massive foundational models, introducing severe computational, latency, and privacy constraints. This paper investigates the operational viability of 7B to 8B parameter Large Language Models (LLMs)—specifically Gemini 3.1 Flash Lite, Llama 3.1 8B, and Qwen 2.5 7B—integrated within the standardized Model Context Protocol (MCP) framework. To evaluate these sub-10B models, we introduce "Complete-MCP," a comprehensive evaluation pipeline engineered to test model efficacy across single-step, sequential, and complex multi-hop tasks in simulated enterprise environments. Experimental results demonstrate that small-to-medium models can achieve a task success rate exceeding 80% when supported by optimized local-first infrastructure, such as adaptive parsing algorithms and heuristic rate-limiting. Llama 3.1 8B showcased exceptional reliability in localized sequential tasks, while Qwen 2.5 7B demonstrated superior logical routing. The findings establish that computationally lightweight autonomous agents can successfully rival high-capability models, offering a scalable, privacy-preserving, and cost-effective blueprint for edge intelligence.
Pages:
5666 - 5672