SaaS · customer support
Support conversation analytics: from chat exports to issues, trends and knowledge-base gaps
A pipeline that turns live-chat exports into product analytics: contact episodes, classified issues, growing topics, incident candidates and knowledge-base gaps, with a deterministic layer that audits the model.
- Industry
- SaaS · customer support
- Technologies
- Python · scikit-learn · rapidfuzz
Challenge
A software company's product, support and documentation teams wanted their support chats to tell them what is broken, what is growing, what customers cannot find and where support time goes. The raw material made that hard. The chat tool keeps one thread per visitor that runs for months, so a thread is not a ticket. One conversation often carries several unrelated problems. And a model that labels thousands of conversations needs a check that does not come from the model itself.
Solution
We built a batch pipeline in Python. Threads are split into contact episodes at 12-hour gaps, a threshold taken from the valley in the distribution of gaps, and speaker roles are recovered with a fixed order of rules. Each episode is broken into issue units, and each unit is classified on 21 fields, including the symptom, the root cause and whether a feature is missing or just hard to find. A regex layer covering about ten languages extracts facts independently of the model and is used to audit it. Clustering, growth detection and percentile-based impact scores then produce the report.
How it connects
- 01 Parse episodes
- 02 Deterministic signals
- 03 LLM classification
- 04 Clustering
- 05 Scoring & audit
- 06 Report & handoff
What we delivered
- Episode segmentation and speaker attribution for long-lived chat threads
- Issue units classified on 21 fields by a fast model, with naming, summaries and audits by a stronger one
- Deterministic signal layer in about ten languages, used as ground truth against the model
- Topic clustering, growth detection and incident candidates: many different customers reporting the same problem within 90 minutes
- Impact and opportunity scores built from percentile-ranked components
- Knowledge-base gap funnel with article briefs for writers, plus churn and competitor signals
- Agent-quality review queue, explicitly not a ranking of people
- HTML report, handoff pages per topic for engineers, and questions in plain language over the whole corpus
Functionality
- Every figure keeps the IDs of its source episodes, so any number opens down to the conversations behind it
- A sample re-classified blind by a different model to measure agreement field by field
- Model calls cached by prompt hash, so an interrupted run resumes without paying twice
- Detection of agent replies saying a feature does not exist, a signal for product and documentation teams
- Unit tests on the deterministic layer
Integrations
- Live-chat conversation exports
- Anthropic Claude models: a fast model for the bulk pass, a stronger one for naming and audit
Technologies
Result
Run on a full month of real support conversations; the report and the engineering handoff pages were delivered. It runs on demand as a batch job, and a dashboard is not part of this phase.
Why keep a rule-based layer next to the model?
Because a model that grades its own work cannot tell you when it is wrong. The regex layer finds facts that can be checked directly, such as an error code, a refund request or an agent saying a feature does not exist, in about ten languages and independently of the model. Where the two disagree, the disagreement is reported rather than smoothed over. That keeps the model useful for the part only it can do, reading meaning, while the numbers stay traceable.
Discuss your project
Tell us about your project