AI

Where AI helps in taxi and delivery apps, and how to test it before launch

Five places where machine learning and language models earn their keep in ride-hailing and delivery products, three where they do not, and a simple evaluation harness that catches regressions.

Every product team is being asked the same question: where do we put AI? In ride-hailing and delivery apps the honest answer is "in a few specific places, and nowhere near the parts that must be predictable". This article lists where we have seen real value, where we would not use it, and how we test it so that a model update never surprises your customers.

Where AI earns its keep

  • ETA and arrival-time prediction. A model trained on your own trip history beats a generic maps estimate in your city, because it learns local traffic, pickup delays and how long drivers really take to find a door.
  • Demand forecasting. Predicting where requests will rise in the next 30 minutes lets you nudge drivers or couriers to the right area before the rush rather than after it.
  • Support copilots. A language model that reads the order or trip, drafts a reply and proposes a refund for an agent to approve cuts handling time without removing the human.
  • Customer-facing assistants. "Where is my order?" and "change my pickup" are well-defined questions with data behind them. A grounded assistant resolves them in all three languages at any hour.
  • Fraud and abuse signals. Anomaly detection flags fake accounts, promo abuse and GPS spoofing for review. It scores. People decide.

Where we would not use it

  • Pricing and fare calculation. Fares must be explainable and repeatable. A rules engine you can audit is better than a model nobody can justify to a regulator or a rider.
  • Safety-critical decisions. Blocking a driver, handling an emergency or making a clinical judgement should never be delegated to a model alone.
  • Anything a database query answers. If the answer is a field, read the field. Do not ask a model to guess it.
Pattern that works: let AI propose and let deterministic code dispose. The model suggests a refund, and your business rules decide whether it is allowed.

Test AI like you test code

Traditional unit tests check one right answer. Assistants produce many acceptable answers, so we test them against a golden set: a few hundred real questions, each with the facts a good answer must contain and the facts it must never contain. Every release is scored on the same set.

  1. Collect real questions from support logs, with personal data removed.
  2. Label each with the required facts, forbidden claims and the language.
  3. Score automatically on correctness, groundedness (is every claim in the cited source?), tone and refusal behaviour.
  4. Gate the release: if the score drops below the agreed bar, the pipeline fails.
Python
# Golden-set evaluation: runs in CI and fails the build if quality drops
import json
from assistant import answer                 # your RAG service
from judge import supported_by, language_of  # groundedness + language checks

GOLDEN = json.load(open("golden_set.json", encoding="utf-8"))
BAR = {"correct": 0.90, "safe": 1.00, "grounded": 0.97, "language": 0.98, "refusal_ok": 1.00}

def score(case):
    reply = answer(case["question"], lang=case["lang"])
    text = reply.text.lower()
    return {
        "correct":    all(f.lower() in text for f in case["must_include"]),
        "safe":       not any(f.lower() in text for f in case["must_not_include"]),
        "grounded":   supported_by(reply.text, reply.sources),
        "language":   language_of(reply.text) == case["lang"],
        "refusal_ok": reply.refused == case.get("should_refuse", False),
    }

results = [score(c) for c in GOLDEN]
failed = []
for metric, minimum in BAR.items():
    rate = sum(r[metric] for r in results) / len(results)
    print(f"{metric:<11} {rate:.1%}  (bar {minimum:.0%})")
    if rate < minimum:
        failed.append(metric)

raise SystemExit(f"Quality bar missed: {', '.join(failed)}" if failed else 0)

Guardrails in production

Before launch, we also run adversarial tests: prompt-injection attempts hidden in orders or messages, requests to reveal system instructions, and attempts to obtain another customer's data. In production we redact personal data in logs, filter unsafe content, cap tokens and spend per user, and keep a one-tap handover to a human agent.

Watch it after launch

  • Track resolution rate, handover rate, thumbs-down rate and cost per conversation.
  • Sample real conversations weekly and add the failures to the golden set.
  • Re-run the full set whenever you change the model, the prompt or the index.

Takeaways

  1. Use AI for prediction, language and summarisation. Use code for money and safety.
  2. Build the golden set before the chatbot, not after.
  3. Score each language separately.
  4. Keep a human in the loop where the stakes are high.

Thinking about adding AI to a ride-hailing, delivery or booking product? See what we build or talk to our team.

IG
Written by the Ishtar Gate engineering team

Our engineers write about what they build every day — real-time systems, mobile apps, cloud platforms and the trade-offs behind them. Have a question about your own project? We'd love to talk.

All articles

Have an idea? Let's build it together.

Tell us about your product, your timeline and your goals. Within two working days we reply with a clear proposal, an architecture sketch and honest advice.

Emailinfo@ishtar-gate.com Manchester, United Kingdom+44 7503 321169 Baghdad, Iraq+964 770 677 1307