I was sceptical at first too, but assuming you have a half-decent/modern GPU and still retained at least one kidney after you bought one in 2026, you should be in a position to deploy your own AI Agent (powered by a Large Language Model running locally) in the comfort of your own home.
Without further ado, let's dive into it.

Link to the Git Hub repository.
Prerequisites
I have used the following for the purpose of this article:
- Memory-savvy GPU (I have used an RX 9070 XT)
- Ollama
- Python 3.10 or newer
- Virtualised networking devices (I have used Containerlab with Arista cEOS images)
It doesn't matter if your devices are virtual or physical; they could be either, as long as SSH works.
If you're going down the Containerlab route, two things caught me out. First, make sure the management subnet in your topology file doesn't overlap with your actual home network, otherwise your host will happily send traffic out of the wrong interface and you'll be wondering why you can't ping anything. Second, if you change the management subnet, destroy the lab with --cleanup before redeploying, as the switches keep their old configs (including the old default route) otherwise.
how does it actually work?
Before we get stuck in, it's worth understanding what an agent actually is, because it's not just the LLM. The model on its own can only take text in and spit text out. It can't SSH to anything.
What makes it an agent is the combination of three things: the model, a loop that keeps the conversation going, and tools, which are just regular Python functions. When the model decides it needs some data, it doesn't run anything itself. It replies with a request along the lines of "please run show version on SW1", our code runs the actual function, and then we send the output back to the model so it can answer the question. The model thinks, our code does. Keep that in mind, as it will make the rest of the article make a lot more sense.
Ollama installation
I will not dive into how to deploy Ollama in much detail. The documentation is pretty comprehensive, and the installation is fairly straightforward. More information can be found here.
For the purpose of this demo, I have used the qwen3:14b language model, which is the most suitable model for my GPU. Feel free to run your current GPU past ChatGPT/Claude to understand which model would be the most suitable for your hardware. Make sure you pull it before you start with ollama pull qwen3:14b (note the colon, not a hyphen, as I found out the hard way).
Python
Let's dive into the fun stuff now.
First things first, create a virtual environment and install the two packages we need. If you're on a recent Debian or Ubuntu, pip will refuse to install packages system-wide anyway, so a venv isn't really optional:
python3 -m venv venv
source venv/bin/activate
pip install ollama netmikoWe will also need the credentials for our switches in environment variables:
export ARISTA_USERNAME=admin
export ARISTA_PASSWORD=admin(If you're on fish, it's set -x ARISTA_USERNAME admin instead.)
I originally started off by going off their docs (link here), and I have just built upon that. Depending on how you got your Ollama set up (or whether you query the LLM as localhost, which is what I have done), your initial draft should look like this:
from ollama import Client
client = Client()
messages = [
{
'role': 'user',
'content': 'Why is the sky blue?',
},
]
for part in client.chat('gpt-oss:120b-cloud', messages=messages, stream=True):
print(part.message.content, end='', flush=True)All I have changed to start with was the model that I am targeting, to qwen3:14b, and then I executed the script just to see if it triggers. Worth noting that the model in their example is a cloud model, meaning it runs on Ollama's servers rather than your GPU, which kind of defeats the point of this article. Assuming your Ollama is listening on localhost and you have downloaded your model, you should be good to go!
But that's not what we want... Let's build some fundamentals for the functions that will be utilised by the Agent once everything is built.
tools.py - our helpers
We know for a fact that, at bare minimum, we need to achieve the following:
- A function that returns the IP address of the queried device based on the device name.
- A function that obtains the credentials. I have used environment variables for the purpose of this article; however, it's completely up to you how you tackle this.
- A function that will execute Netmiko, query the devices with the show commands and return the output.
Below is what I have drafted to tick all of the above nicely for us.
import os
from netmiko import ConnectHandler
def obtain_credentials():
username = os.environ["ARISTA_USERNAME"]
password = os.environ["ARISTA_PASSWORD"]
return username, password
def get_single_device_addr(host):
DEVICES = {
"SW1" : "10.0.101.11",
"SW2" : "10.0.101.12",
}
return DEVICES[host]
def run_show_command(host, command):
username, password = obtain_credentials()
ip = get_single_device_addr(host)
net_connect = ConnectHandler(
device_type="arista_eos",
host=ip,
username=username,
password=password,
)
output = net_connect.send_command(
command
)
net_connect.disconnect()
return outputI am not going to go through the code too much, but the plan is that the run_show_command function is what the agent will be using to query the devices. As you can see, the function already handles the credentials; therefore, we won't have to pass any passwords or sensitive data to the LLM.
It's worth testing this on its own before any AI gets involved. Import it in a Python shell and run run_show_command("SW1", "show version"). If you get output back, you know your Netmiko side works, and anything that breaks later is on the model side.
main.py - the brain of the operation
getting started
Now, I appreciate that if you have done any Network Automation with Python in the past, the above is relatively simple, so let's take a bit more time and explain the concept of what we will be doing in main.py.
Let's modify the code that we were given from their documentation slightly.
I will do my best to split the code into sections so that it makes more sense; we will also implement the full 'solution' step by step. Now, what do we need?
Below is a single-prompt example without the while loop that handles the back-and-forth with the agent.
from ollama import Client
from tools import run_show_command
client = Client()
messages = [
{
'role': 'system',
'content': "You're a network assistant for a lab with two Arista switches called SW1 and SW2, you can run read only show commands and if you don't know something, say so."
},
{
'role': 'user',
'content': 'what is the firmware version of SW1 and SW2',
},
]
chat = client.chat(
'qwen3:14b',
messages=messages,
tools=[run_show_command],
think=False,
)
reply = chat.message
print(reply)import the function that we created
from tools import run_show_command allows us to import the function that we created in tools.py, which SSHes to the device and runs the show command, along with the credentials retrieval.
messages variable (list)
messages = [
{
'role': 'system',
'content': "You're a network assistant for a lab with two Arista switches called SW1 and SW2, you can run read only show commands and if you don't know something, say so."
},
{
'role': 'user',
'content': 'what is the firmware version of SW1 and SW2',
},
]As you can see, we have diverged a bit from the original concept and created two separate roles, system and user. What are roles, you might ask? We distinguish the following message roles:
- system: Sets the behaviour, context, rules, or persona for the model. It is usually the first message in a conversation.
- user: Represents input or prompts sent by the person interacting with the model.
- assistant: Represents responses generated by the model.
- tool: Used to feed the results of an executed function or tool back into the conversation history so the model can read and process them.
Based on the above, the concept of roles becomes quite straightforward, assuming that you haven't lived under a rock and have used an AI agent of any kind before. In essence, the core identity is retained in the system role, and the user role represents our regular prompt, the same as what we would normally type into a webchat through ChatGPT or Claude or whatever.
client.chat
chat = client.chat(
'qwen3:14b',
messages=messages,
tools=[run_show_command],
think=False,
)
reply = chat.message
print(reply)Now, let's unpack the client.chat function call. We have included the following:
qwen3:14b, which is our model that was pulledmessages=messages, where we associate the messages listtools, where we associate our external functions. Note that we're passing the function itself, without brackets. We're not running it here, just letting the model know it exists.think=Falseis optional. It switches off the model's reasoning step (the thinking you'll see below), which makes it noticeably faster. The trade-off is that on more complex, multi-step questions it can be less accurate, so it's worth playing with both.
At the very end, I have allocated the actual response into a variable called reply.
Now it looks like we are good to go, no? Let's give it a test!
venv ❯ python main.py
role='assistant' content='' thinking=None images=None tool_name=None tool_calls=[ToolCall(function=Function(name='run_show_command', arguments={'host': 'SW1', 'command': 'show version'})), ToolCall(function=Function(name='run_show_command', arguments={'host': 'SW2', 'command': 'show version'}))]Just to show you the difference that think=False makes, here is the output with True instead:
role='assistant' content='' thinking="Okay, the user is asking for the firmware versions of SW1 and SW2. Let me think about how to get that information.\n\nFirst, I remember that on Arista switches, the firmware version can be found using the 'show version' command. This command displays detailed information about the switch, including the software and firmware versions.\n\nSo, I need to run 'show version' on both SW1 and SW2. The user mentioned that I can run read-only show commands, so this should be possible. \n\nWait, but the tools provided have a function called run_show_command which takes a host and a command. The host is the switch name, like SW1 or SW2, and the command is the CLI command to execute. \n\nTherefore, I should make two separate function calls: one for SW1 with 'show version' and another for SW2 with the same command. Then, the output from each command will include the firmware version. \n\nI should check if there's anyother command that might give firmware info, but I think 'show version' is the standard one. Arista's EOS (Extensible Operating System) uses this command extensively for system information. \n\nSo, the plan is to call run_show_command for each switch, parse the output to find the firmware version, and then present both versions to the user. If there's an error in retrieving the information, I should inform the user accordingly. \n\nBut since the user hasn't provided any errors yet, I'll proceed with making the two function calls.\n" images=None tool_name=None tool_calls=[ToolCall(function=Function(name='run_show_command', arguments={'host': 'SW1', 'command': 'show version'})), ToolCall(function=Function(name='run_show_command', arguments={'host': 'SW2','command': 'show version'}))]tool calls
Here is the interesting part. The response is the model's own message, and as you can see, its content is empty. Instead, it contains what's called tool_calls.
In essence, the way to interpret that is: if an agent's response includes tool_calls, it means that it is not done yet. It's telling us which function it wants to run and with which arguments, based on the prompt we have given it. Notice that nothing has actually touched the switches yet; the model has only asked. Unfortunately, our current code doesn't do anything with that request, since there are no conditional statements or loops in our script; therefore, the conversation ends here, as is.
Here is what we need to do, and it is quite simple, to be fair. We need a while loop that only breaks when tool_calls is None, which basically means that the agent has answered our question and there is nothing more it needs to run to answer our prompt.
an example of a working loop
while True:
chat = client.chat(
'qwen3:14b',
messages=messages,
tools=[run_show_command],
think=False,
)
reply = chat.message
if reply.tool_calls is not None:
messages.append(reply)
tool_calls = reply.tool_calls
for tool in tool_calls:
tool_name = tool.function.name
args = tool.function.arguments
try:
funct_name = dispatch_dictionary.get(tool_name, "n/a")
output = funct_name(**args)
message = {
"role": "tool",
"content": output,
"tool_name": tool_name,
}
messages.append(message)
except Exception as e:
print(f"[ERROR] : {e}")
else:
print("No further (or any) tool calls found.")
print(reply.content)
breakLet's break the code above into smaller sections and validate what each part handles so that it makes more sense!
breakdown
if reply.tool_calls is not None:
messages.append(reply)
tool_calls = reply.tool_callsAs mentioned above, we need to check whether tool_calls is not None, as this means the agent still isn't done. This is why it is important to append the reply to the messages list, as it records the model's own request: "please run show version on SW1". Without it, the history has no record that the model asked for anything, and remember, the model has no memory of its own. The messages list is all it's got.
I have also allocated the tool calls to a variable, which is self-explanatory.
for tool in tool_calls:
tool_name = tool.function.name
args = tool.function.arguments
try:
funct_name = dispatch_dictionary.get(tool_name, "n/a")
output = funct_name(**args)
message = {
"role": "tool",
"content": output,
"tool_name": tool_name,
}
messages.append(message)
except Exception as e:
print(f"[ERROR] : {e}")In the above, we are looping through each tool call and unpacking it as we go. Here is an example of what both tool_name and args look like:
tool_name = run_show_command
args = {'host': 'SW1', 'command': 'show version'}We are unpacking tool_name and args for one reason: the model has only given us the name of the function as a piece of text, so we need to turn that back into the actual function and run it ourselves. To do that, we rely on a simple dictionary that maps the tool_name to the real function (this pattern is usually called a dispatch table, if you want to read up on it).
dispatch_dictionary = {
"run_show_command" : run_show_command,
}As you can see, funct_name returns the run_show_command function that is imported at the top of the script. Now that we have the function allocated to funct_name, all we need to do is follow the run_show_command(host, command) formula that we built in tools.py by passing **args, and allocate the result to its own variable called output. The ** simply unpacks the dictionary into keyword arguments, so it's the equivalent of writing run_show_command(host="SW1", command="show version") by hand.
the continuation of the message variable
tool: Used to feed the results of an executed function or tool back into the conversation history so the model can read and process them.
Considering that we are now dealing with a different type of role called tool, we need to wrap the output in a message and append it to messages. That way, we retain the history of the current conversation, so the agent knows what we asked, what it requested, and what came back. The loop then goes round again, and this time the model has everything it needs to actually answer.
Last but not least, should reply.tool_calls be None, let's just print the content of the response and call it a day! Remember, no more tool calls means the Agent has done its job (oh, and don't forget to break out of the while loop!).
python
else:
print("No further (or any) tool calls found.")
print(reply.content)
breakFull main.py
python
from ollama import Client
from tools import run_show_command
client = Client()
messages = [
{
'role': 'system',
'content': "You're a network assistant for a lab with two Arista switches called SW1 and SW2, you can run read only show commands and if you don't know something, say so."
},
{
'role': 'user',
'content': 'what is the firmware version of SW1 and SW2',
},
]
dispatch_dictionary = {
"run_show_command" : run_show_command,
}
while True:
chat = client.chat(
'qwen3:14b',
messages=messages,
tools=[run_show_command],
think=False,
)
reply = chat.message
if reply.tool_calls is not None:
messages.append(reply)
tool_calls = reply.tool_calls
for tool in tool_calls:
tool_name = tool.function.name
args = tool.function.arguments
print(f"[DEBUG]: NAME = {tool_name} \n ARGS: {args} \n DISPATCH_DICT = {dispatch_dictionary}")
try:
funct_name = dispatch_dictionary.get(tool_name, "n/a")
output = funct_name(**args)
message = {
"role": "tool",
"content": output,
"tool_name": tool_name,
}
messages.append(message)
except Exception as e:
print(f"[ERROR] : {e}")
else:
print("No further (or any) tool calls found.")
print(reply.content)
breakwhat happens now?
Let's run the script and see what happens:
venv ❯ python main.py
The firmware version (software image version) for both SW1 and SW2 is:
- **SW1**: `4.32.0F-36401836.4320F (engineering build)`
- **SW2**: `4.32.0F-36401836.4320F (engineering build)`
Both switches are running the same firmware version.As you can see, we have successfully run a show command on an Arista device without overcomplicating the function calls. We are literally using one simple, agnostic function (well, agnostic enough for the Arista platform, but that can easily be changed).
things that caught me out
A few things that tripped me up along the way, so hopefully they won't trip you up:
- The content of a tool message has to be a string. Our run_show_command returns a string, so we're fine here, but if you add a function that returns a dictionary (an inventory, for example), you'll get a Pydantic validation error. Wrap the output in
json.dumps()and you're sorted. - Watch the indentation of your else. If it ends up lined up with the for loop instead of the if statement, Python treats it as a for/else, which runs as soon as the for loop finishes. The script then breaks straight after running the tools and never goes back to the model for an answer.
format="json"doesn't play nicely with tool calls. I tried forcing the final answer into JSON using the format argument, and the model ended up writing its tool request as JSON text instead of actually calling the tool. If you want JSON out, let the loop finish first and then make one extra call with format set and no tools.
what happens later?
Well, that's the fun part, isn't it? You'll notice I have deliberately kept this read-only. We could now easily add set command functions, but those deserve proper safety guards before executing them, and a line in the system prompt saying "please be careful" isn't one of them. Dynamic inventory? Chatbot? Logging to retain history? The options are endless.
