Quick Answer
An AI customer support agent is a language model that can call small functions in your code. The functions look up an order or search your return policy, and the model writes a reply based on the results. A basic agent for an online store needs a chat model with function calling, an embedding model for policy search, and about 170 lines of Python. It answers order-status and policy questions on its own and routes refunds and complaints to a person.
Introduction
Most online stores answer the same few questions all day. Where is my order? How long does a refund take? Can I change the delivery address? A human agent can answer them, but every answer still costs a few minutes of reading, searching, and typing.
An AI agent does that reading and searching itself. It does not only write text. It looks up the order in your system, checks the store policy, and replies with facts from your own data. This article shows which requests an agent can close without a person and how an agent is built. Then it walks you through writing a working one in Python, step by step. You can copy every command and code block.
Terms in Plain Words
| Term | What it means |
| API key | A long password that lets your program use a service. You create it once in your account. |
| Terminal | A window where you type commands. On Windows, it is PowerShell; on a Mac, it is the Terminal app. |
| Function calling | The model answers with the name of a function it wants to run instead of text. Your code runs the function. |
| Tool | A function in your code that the model is allowed to ask for. |
| Embedding | A list of numbers that describes a text’s meaning. Similar texts get similar numbers. |
| Knowledge base | The texts the agent may quote, such as return rules. |

Which Requests an Agent Can Close on Its Own
An agent works well when the answer already exists in a system it can read. Order status sits in the order database. Return rules sit on the policy page. Both can be fetched in a second, so the agent does not need to guess.
It works badly when the request needs a decision. A refund outside the return window, a damaged item that needs a photo check, or a customer who asks for a manager all belong to a person. In those cases, the agent should pass the conversation on with the order number and the question already collected so the human doesn’t ask for them again.
| Request | What the agent reads | Who closes it |
| Where is my order? | Order system and carrier tracking | Agent |
| How long do refunds take? | Refund policy text | Agent |
| Can I return this item? | Return policy and order date | Agent |
| Refund outside the return window | Order history and policy exceptions | Human |
| Damaged item | Photos and order history | Human |
| Complaint about a staff member | Conversation history | Human |
A request goes to a person when neither the order data nor the policy text contains the answer.
What an Agent is Made of
A support agent has four parts. The language model reads the customer message and decides what to do next. Tools are functions in your code that read data or perform an action. A knowledge base holds the text the agent can quote, such as return rules. A loop calls the model again after each tool result until the model has a final answer.
The model never runs a tool itself. It returns the name of a function and its arguments, and your code runs the function. OpenAI describes this flow in its function calling guide as five steps. You send the request with a list of tools. The model returns a tool call. Your application executes it. You send the output back. The model then answers or asks for another call.

The customer message goes to your code, your code asks the model, and the model answers or names a tool.
This arrangement also protects the store. Because your code runs every tool, you decide which functions exist, which arguments they accept, and which customer’s data they may touch.
The knowledge base is searched by meaning rather than keywords. An embedding model turns each policy paragraph into a list of numbers called a vector. The customer question is turned into a vector in the same way, and the paragraphs with the closest vectors are returned. A customer who writes “how fast do I get my money back” still finds the refund paragraph, even though the question is missing the word refund.
Which Models the Agent Needs
A support agent uses up to three model types. A chat model with function calling runs the loop. An embedding model indexes the knowledge base and searches it. Speech models are needed only if customers call or send voice messages. Speech-to-text transcribes the caller, and text-to-speech reads the answer aloud.
You can use separate vendors, but then you keep several accounts, keys, and invoices. A gateway such as AI/ML API puts them behind one OpenAI-compatible endpoint: https://api.aimlapi.com/v1. According to its documentation, function calling is available for models from Anthropic, OpenAI, Google, Alibaba, Meta, DeepSeek, Mistral, and xAI, among others.
The platform also lists the embedding models text-embedding-3-small and text-embedding-3-large, speech-to-text from Assembly AI, Deepgram, and OpenAI, and text-to-speech from ElevenLabs, Deepgram, and Microsoft.
This is model access only. A gateway does not give you a finished support agent. The tools, the knowledge base, and the loop stay in your code, and the next section builds them. One practical result is that the agent code does not depend on the model. To find out whether a cheaper model handles your tickets well enough, you change one string.
Set up Your Computer

You do this once. It takes a few minutes.
Install Python
Download Python 3.10 or newer from python.org and install it. To check the installation, open the terminal and run python –version. On a Mac, use python3 –version. You should see a version number.
Get Your API key
Create an API key in your AI/ML API account and copy it. Treat the key like a password. Don’t paste it into the code file or send it to anyone.
Create the Project Folder
Open the terminal and run these two commands. The first creates a folder, and the second moves into it.
Install the OpenAI Package
The package provides the OpenAI client the code uses. On a Mac or Linux, run these commands.
On Windows PowerShell, run these commands instead.
If PowerShell says that running scripts is disabled, run Set-ExecutionPolicy -Scope Process RemoteSigned and activate the environment again. After activation, the terminal line starts with (.venv).
Save Your Key in an Environment Variable
An environment variable keeps the key outside your code. On a Mac or Linux, run this command with your own key.
In Windows PowerShell, run this.
The variable exists only in this terminal window. If you close the window, set it again.
Building the Agent in Python

Create a file named support_agent.py in the support-agent folder. Add the four code blocks below to it one after another, in the order shown. The example uses the official OpenAI Python package and a store with three tools. The policies and orders are sample data. In a real store, you replace them with calls to your own systems.
Step 1. Connect the Client and Index the Policies
The first block connects to the API, stores the sample data, and defines the policy search. The vectors for the policies are created the first time a customer asks a question and are reused after that. In production, store them in a database. Cosine similarity measures how close two vectors are, and search_policies returns the two closest paragraphs.
Step 2. Write the Tools
The model can fill in only order_id. The customer_id comes from your login session and is added by the code in step 4. A customer cannot read another person’s order, even if they ask the agent to look it up. The extra customer_id argument on the other two functions only gives all handlers the same call shape.
Step 3. Describe the Tools to the Model
The model picks a tool by reading its description, so write each description as a plain instruction. The system prompt sets the boundary. Policy answers must come from the search result, and anything the tools cannot answer goes to a person.
Step 4. Run the Loop
Each pass sends the whole conversation to the model. If the reply has no tool calls, it is the final answer. Otherwise, the code runs every requested tool and appends each result as a message with the role tool and the matching tool_call_id. A tool that raises an error returns the error text as its result, so the model can explain the problem or hand over instead of crashing the request. The max_steps limit stops a conversation that keeps calling tools without finishing, and the agent then hands over to a person.
With the sample question, the model should call get_order_status and search_policies in the same turn and write one answer from both results. The reply wording will differ between runs. The loop follows the same pattern as Anthropic’s tool_use documentation: the model stops with stop_reason set to tool_use, and you reply with a tool_result block. If you later call Claude through its own API, only the message format changes.
Step 5. Check the Tools without the Model
Before you run the whole agent, check that your own tools work. Create a second file named check_tools.py in the same folder.
Run it with python check_tools.py. This script does not call the AI model.
Terminal output of check_tools.py showing an order status, a refusal for another customer, and a handoff line
You should see three lines. The first is the order for customer c42. The second shows what happens when customer c77 asks for the same order. The tool refuses, as described in step 2. The third line is the handoff message.
Step 6. Run the Agent
Now run the whole agent.
The script asks, “Where is order A1001 and how long do refunds take?” as customer c42. After a few seconds, it prints one answer. The wording is different on every run. The answer should say that order A1001 was shipped with DHL and is expected on October 9, 2026. It should also say that refunds take 5 business days after the warehouse receives the item. If you see this, the agent works.
To try other questions, change the text in the last line of the file. These three cases show the main behaviors.
| Question | What should happen |
| Where is order A1002? | The agent says it cannot find this order for the customer because A1002 belongs to another customer. |
| I want a refund for order A1001 | The agent calls handoff_to_human and the terminal prints a line that starts with [handoff]. |
| Can I return a used item? | The agent answers from the return policy and says that the item must be unused. |
If Something Goes Wrong
| Message in the terminal | Cause | Fix |
| KeyError: ‘AIMLAPI_KEY’ | The key variable is not set in this terminal window | Run the export or $env command again |
| ModuleNotFoundError: No module named ‘openai’ | The environment is not active, or the package is not installed | Activate the environment and run pip install openai |
| An authentication error or status 401 | The key is wrong or has extra spaces | Copy the key again and set the variable again |
Where Agents Fail and How to Limit the Damage
The most common failure is a confidently wrong answer. An agent that answers a return question from memory instead of from the policy text can invent a 60-day return window. The system prompt in the example forbids this, but a prompt is not a guarantee. Log every tool call and every final answer, and have someone read a sample each week.
The second failure is over-permissioning. Start with read-only tools. Issuing a refund, changing an address, or applying a discount should require human approval until the logs show the agent requests them correctly.
Customer data needs care as well. Send the model only the fields it needs to answer. The order tool above removes the customer ID before it returns an order, and a real system should also leave out the full address and any payment details.
Test before launch. Take 50 to 100 past tickets, run them through the agent, and ask a support employee to mark each answer as correct, wrong, or needing a human. Repeat the test whenever you change the prompt or the model. The model is a single string in the code, so the same test set also shows whether a cheaper model is good enough.
How to Measure whether the Agent Helps
Track a small set of numbers from the first week.
| Metric | How to count it | What to watch |
| Resolution rate | Conversations that ended without a handoff, divided by all conversations | A rise, together with more reopened tickets, means customers gave up |
| Handoff rate | Conversations where the agent called handoff_to_human | Group the reasons and add tools for the common ones |
| Reopen rate | Customers who write again about the same order within a few days | Wrong or incomplete answers |
| First response time | Time from the customer message to the first reply | Automated answers should arrive within seconds |
| Cost per conversation | Model cost of all calls in one conversation | Compare models on the same test set |
Read the handoff reasons first. They show which requests the agent cannot handle yet. If many handoffs ask where a parcel is after the delivery date has passed, add a carrier tracking tool. The agent grows one tool at a time, and each new tool is justified by the logs.
FAQs

What is an AI customer support agent?
It is a program in which a language model reads a customer message, calls functions that fetch real data, and writes a reply from the results. In an online store, it can check an order, find a return rule, and pass the chat to a person when it cannot answer.
How is an AI agent different from a chatbot?
A scripted chatbot follows fixed buttons and keywords. An AI agent understands free text and can call tools, so it answers with data from your systems. A language model without tools only writes text and can invent facts about orders or policies.
What do I need to build an AI support agent?
You need a chat model that supports function calling, an embedding model for searching your policies, access to your order data, and a short Python program. The example in this article has about 170 lines and uses sample data.
Which requests should an AI agent not handle?
Anything that needs a decision. Refunds outside the return window, damaged items, payment disputes, and staff complaints should go to a person. Start with read-only questions such as order status and policy rules.
How do I stop the agent from making up answers?
Give it tools for facts and tell it in the system prompt to answer only from tool results. Log every answer and have a person read a sample each week. Test the agent on 50 to 100 past tickets before launch and after every prompt or model change.
Can I use a different model than gpt-4o?
Yes. The model name is one string in the code, in the CHAT_MODEL line. Use any chat model that supports function calling, then rerun your test tickets after the change.
How do I know the agent saves time?
Track the resolution rate, handoff rate, reopen rate, first response time, and cost per conversation. Read the handoff reasons first, because they show which requests the agent cannot handle yet.
Conclusion

An AI agent takes over the requests whose answers already exist in your systems. It works when its tools are limited, its answers come from your own text, and a person receives everything the agent cannot decide. The Python example above is short, but it contains the parts a production agent needs. Start with order status and policy questions, study the handoff reasons for a few weeks, and add tools only where the numbers show a gap.



















