Three pieces of work I can talk about in detail. The first is the production chatbot I ran at Standard Chartered. The other two are prototypes I built in 2026, with public code. For each one I’ve written down the problem, my role and the trade-offs I made. The first two also cover how I tested them.
Running a bank’s conversational AI in production (Standard Chartered, 2018 to 2021)
Problem
Standard Chartered Hong Kong ran Stacy, its enterprise virtual assistant, on the public website, online banking and the mobile app. It used traditional NLP models to answer general banking questions, and it connected to backend services for things like finding a nearby ATM or branch. The main business goal was cost saving. In a bank, giving one customer another customer’s balance is a failure nobody can afford, so the assistant had to answer correctly and hand off what it couldn’t handle.
My role
Chatbot Specialist, then Associate on the chatbot team. I coordinated the HKD 5M+ programme from pre-launch planning through launch and post-launch optimisation, and I was the operating owner in production. The work covered NLP improvement, model-training processes, vendor coordination, production releases and digital measurement.
Trade-offs
- Containment vs handing off: questions below a confidence threshold went to live agents.
- What to automate in testing: FAQ-style intents returned the same wording every time, so I automated those checks with exact match (seed question in, expected answer out). Answers tied to a backend and a real customer, such as an account balance or an ATM location, were still tested by a person.
Evaluation
- Exact-match regression runs for FAQ intents, which cut the time spent clicking through the chatbot one prompt at a time.
- Manual testing for backend-linked answers.
- Regular review of real customer conversations across the FAQ intents, with annotation. Customers didn’t always phrase things the way the training data expected, and the review fed back into model and journey improvement.
- On the measurement side, we used Adobe Analytics to track which FAQs people asked before they reached a product application page, and reported that monthly in Power BI.
Result
About 80% automated containment across the public website, online banking and the mobile app.
Insights from the conversation review and annotation contributed to reducing the chatbot’s negative-feedback rate from 7.4% to 4.1% and lifting year-over-year cost savings by 27.5%.
Related writing: Exact match used to QA a bank chatbot. It cannot QA an LLM., A chatbot question is a conversion touchpoint. Treat it like one. and In the AI Era, the Data Signal Still Comes First
Documentation RAG assistant for GA4 and GTM (2026, prototype)
Problem
Non-technical marketers ask AI how to track something in GA4 or Google Tag Manager. They get a fluent answer and no easy way to check it. The answer might mix up Universal Analytics and GA4, or recommend a developer task that GA4 already handles. I wanted to test whether answers grounded in Google’s official documentation, with the source pages shown, would be easier to check and trust.
My role
I built it on my own: crawl, storage, indexing, retrieval, prompts, the evaluation harness and the UI.
What I built
A local Python app. It crawls official GA4 and GTM help pages into Notion (about 3,199 pages visited and 3,178 stored in the last crawl), removes duplicates, then indexes them in Chroma using LangChain markdown chunking. Chat and evaluation run in a Gradio app, and the retrieved pages sit beside each answer.
Trade-offs
- Parent-child chunking: short passages (about 500 characters) are embedded for search, and the full section (up to about 2,800 characters) goes to the model.
- I dropped the LLM query rewrite, expansion and rerank from live retrieval. They were slower and still surfaced the wrong pages (ads, API, generic setup). A simple follow-up merge, a few GA4/GTM query variants and a hybrid score (vector plus lexical overlap on titles, URLs and headings) worked better.
- Google’s help pages repeat the same setup intro, so I filtered out intros that had no procedure words. Otherwise every question retrieved setup boilerplate.
- Vague questions get clarifying questions first. If the docs are thin, the assistant says so and sends the marketer to a person who knows GA4 and GTM.
- I used LLM-as-judge instead of human review for every answer, because I needed to re-run the evaluation after each change. I treat it as a first filter. It can’t sign off on its own.
Evaluation
An LLM judge (GLM) scored each answer from 0 to 1 on accuracy, completeness and relevance, using 10 questions generated from the index. In that initial evaluation the assistant achieved mean scores of 0.945 for accuracy, 0.885 for completeness and 1.000 for relevance. The test was limited, and it wasn’t a controlled comparison against an ungrounded chatbot or a user test with marketers and developers.
Result
Answers stayed on topic and cited the pages they used. The most important miss: asked about the Campaign Manager 360 report’s attribution model, the answer named non-direct last click correctly, then inverted the rule in its plain-language summary. The citation and the judge both made the slip visible, and it’s the reason I keep a human on attribution-type questions. Next steps I’d take: a hand-written gold set built on real briefing questions, and indexing written practitioner notes next to the official docs. The test set leaned toward reporting questions, so I read these scores as an early signal and nothing more.
Code: github.com/anthony-tsui/marketing_tracking_poc · Demo: youtu.be/FourfeZ4fGM · Write-up: Asking AI about GA4 is easy. Trusting the answer is the hard part.
Multi-agent tracking proposal reviewer (2026, prototype)
Problem
Before a marketing team adds tracking to a website, someone has to check the brief against what the site already has. MarTech, developers, compliance and the business each look at it differently, and that review is slow to coordinate.
My role
I built it on my own as one of the five projects from my career break. It’s a prototype, not a live client system.
What I built
You give it a marketing brief and a live URL. A Selenium scraper inspects the site’s current tagging: GTM container IDs, GA4 measurement IDs, third-party tagging scripts, the dataLayer, and GA4 ecommerce events seen during a simulated product, cart and checkout flow. Then four LLM agents work on it:
- MarTech goes first and proposes a dataLayer schema and GTM tags, triggers and variables
- Developer and Compliance each review that proposal side by side. Developer looks at technical feasibility and implementation risk. Compliance looks at privacy and Hong Kong compliance concerns
- Business reads all three and gives a go, no-go or conditional-go recommendation
It runs in a notebook with a local Gradio UI, through an OpenAI-compatible client with a tool-calling loop.
Trade-offs
- If Selenium can’t run, it falls back to static HTML. That’s fine for a first look, but it can’t validate runtime JavaScript, consent behaviour, click events or dataLayer changes.
- Simulated clicks are heuristic. Checkouts that need a login, a region choice or anti-bot handling may not be testable.
- I used separate agents because I wanted to try different perspectives on the same brief.
- The compliance output is LLM-assisted and isn’t legal advice. The prototype doesn’t store results or keep a production-grade audit log.
Code: github.com/anthony-tsui/tracking_proposal_poc
Also public, from the same career-break set: MMM data-prep code generator, synthetic MMM data generator, competitor-analysis pipeline and an advertising and analytics tag auditor (github.com/anthony-tsui).