AI for D2C

10 Questions to Ask Any AI Agent Vendor Before You Sign

A practical buyer checklist for AI agents: action permissions, hallucinations, data isolation, evals, human handoff, observability, pricing and exit terms.

Lokesh Sharma·August 18, 2026

AI agents are easy to demo and hard to evaluate

A polished demo can make almost every AI agent look capable. The scripted question works, the response is fast, the product knows the right policy and the happy path ends cleanly.

Production is different. Customers ask incomplete questions. Policies conflict. Catalog data goes stale. APIs fail. A model is uncertain but still sounds confident. One action has a ₹200 downside and another has a ₹20,000 downside.

That means the real buying decision is not "Can the model answer this demo question?" It is "Can this system behave safely, measurably and economically when the happy path breaks?"

These ten questions are designed for that evaluation. Ask them of any AI agent vendor, including Eldor.

1. What can the agent actually do, not just answer?

Separate conversational capability from action capability. Can the agent only draft responses, or can it search products, modify orders, create tickets, issue refunds, apply discounts, update CRM fields or trigger outbound messages?

Ask the vendor to list every production action the agent can take and the systems it can touch. Then ask which actions are read-only, reversible, approval-gated or fully autonomous.

Good answer: a clear capability map with tool-level permissions.

Red flag: "The agent can do whatever your team can do" without a permission model.

2. What happens when the agent is uncertain or the knowledge is missing?

Every production agent will encounter questions it cannot answer reliably. The important part is whether uncertainty changes behaviour.

Ask what happens when the system has no source, conflicting sources, low confidence, stale data or an unavailable tool. A mature system should have explicit fallback behaviour such as asking a clarifying question, refusing to guess, escalating to a human or limiting the action it can take.

Good answer: defined fallback paths plus examples of when they trigger.

Red flag: confidence is treated as a tone problem instead of a control problem.

3. What are the permission boundaries for high-impact actions?

AI agents become materially riskier when they can take actions. Refunds, discounts, cancellations, address changes, outbound messages and account modifications should not all share the same autonomy level.

Ask whether permissions can be limited by role, amount, order state, customer type, channel or tool. For higher-risk actions, ask whether a human approval step can be required before execution.

Current agent-security guidance emphasizes least-privilege permissions and human approval for high-risk or irreversible actions. Your vendor should be able to explain how those principles show up in the product.

Good answer: "The agent can refund up to X under these conditions, but anything above that requires approval."

Red flag: one global on/off switch for autonomy.

4. How is my company, customer and conversation data isolated?

For a multi-tenant product, ask about isolation at every layer that matters: database access, vector or search indexes, caches, agent memory, logs and analytics.

Also ask whether your data is used to train or improve models shared across customers, whether that behaviour is opt-in or opt-out and what third-party model providers receive.

Do not stop at "your data is encrypted". Encryption is important, but it does not answer whether one agent can retrieve or reason over another customer's data.

Good answer: explicit tenant boundaries, retention controls and a clear model-training policy.

Red flag: vague statements such as "enterprise-grade security" with no architecture or policy detail.

5. How does the agent know which source is current and trustworthy?

Agents often draw from multiple sources: website pages, product catalog, policy documents, order systems, CRM data and previous conversations. These sources can disagree.

Ask how the system handles freshness, source priority, document versions and conflicting information. If the returns policy says 7 days in one document and 14 days in another, which one wins and can an operator see why?

Good answer: source hierarchy, timestamps, sync behaviour and inspectable knowledge.

Red flag: "We crawl your website and the AI figures it out."

6. What evals do you run before and after deployment?

Ask for more than model accuracy. A production agent should be evaluated on the tasks that matter to your business.

Useful eval dimensions can include answer correctness, groundedness, tool-call success, policy adherence, escalation quality, task completion, unsafe-action attempts, language handling and latency.

Then ask the operational question: what happens when the prompt, model, tool or knowledge base changes? Is the same evaluation set rerun before release?

Good answer: named evaluation metrics, sample test cases, failure thresholds and regression testing.

Red flag: evaluation means manually trying ten prompts after every change.

7. How does human handoff actually work?

"We support human handoff" can mean anything from creating a ticket to transferring the entire conversation with useful context.

Ask what triggers a handoff, where the conversation appears, what customer and order context is transferred, whether the human can see what the agent already tried, and whether the agent can resume after the human is done.

Also ask how urgent cases are prioritized. A ₹30,000 purchase blocked by a delivery question should not necessarily wait behind a low-urgency admin request.

Good answer: explicit triggers, full context transfer, ownership and a clear resume path.

Red flag: "We open a Zendesk ticket" with no continuity.

8. Can I audit why the agent answered or acted the way it did?

When something goes wrong, you need more than the final chat transcript. Ask whether operators can inspect relevant sources, tool calls, actions, failures, approvals, handoffs and timestamps.

You do not need access to hidden model reasoning. You do need enough operational trace to answer practical questions such as: Which policy did it use? Which order did it modify? Which tool failed? Who approved the refund?

Good answer: structured logs and audit trails designed for operations, not only engineering.

Red flag: the only debugging interface is the customer conversation.

9. What does the pricing model do during my busiest month?

AI-agent pricing can be based on seats, conversations, resolutions, model usage, messages, automation actions or a mixture of these. The unit matters less than whether the bill is predictable.

Give the vendor a real high-traffic scenario. For example: double your normal conversations, more WhatsApp volume, longer festive threads and a higher share of complex support cases. Ask them to model the bill.

Also ask what counts as a conversation or resolution, whether failed or escalated conversations are billed, which model tiers are included and whether channel costs sit on top.

Good answer: a transparent calculator or scenario-based quote.

Red flag: a cheap headline price with undefined usage units.

10. What happens to my data, workflows and learnings if I leave?

Exit terms reveal how much control you really have. Ask what can be exported: conversations, customer labels, knowledge documents, analytics, workflow configuration and agent settings.

Then ask about deletion timelines, backups, third-party processors and whether any of your data remains in shared training or evaluation datasets.

Good answer: documented export options and deletion commitments.

Red flag: your operating history is effectively trapped inside the vendor.

A simple 20-point vendor scorecard

Score each question from 0 to 2:

  • 0: vague, unavailable or based on trust.

  • 1: partially supported, manual or limited.

  • 2: specific, configurable, observable and demonstrated.

A vendor can still be the right choice with a lower score if the use case is narrow and low-risk. The purpose of the score is not to create a universal cutoff. It is to make tradeoffs visible before deployment.

The pattern behind all ten questions

Every question is really asking the same thing: what happens at the edges?

The happy path is easy to stage. Production quality shows up when the model does not know, the tool fails, the data conflicts, the customer changes channel, an action carries financial risk or a human needs to take over.

The best AI-agent vendors should be comfortable showing those cases, not hiding them.

At Eldor, this is also how we think about commerce agents: one shared customer and store context, explicit actions and guardrails, measurable outcomes, and human control where it matters.

Want more like this?

founders@eldor.ai and we'll add you to the list.