Marketing · social listening

Community listening engine: buying intent and brand mentions from public Telegram groups

A listening engine that scans public Telegram communities, finds buying-intent messages and brand mentions, and sends them to a sales team as explained cards with a draft reply, using a language model only where it is needed.

Industry
Marketing · social listening
Technologies
Python 3.12 · Telethon · PostgreSQL + pgvector
ABC

Challenge

A marketing and sales team wanted to find people in public Telegram communities who are asking for a product like theirs, and to see brand mentions as they happen. Telegram makes this hard in three ways. Accounts are rate-limited and penalized for aggressive reading. Messages are full of duplicates, forwards, look-alike characters and mixed languages. And running every message through a language model would cost far more than the leads are worth.

Solution

We built the engine in Python on a Telegram client layer with an account pool, per-account operation budgets and flood-wait handling. New sources are found through links, forwards and web catalogs and pass a qualification funnel. Messages are normalized, deduplicated at four levels and matched with an Aho–Corasick automaton, with a semantic second pass on pgvector. An intent-by-domain gate lets the model run only when a message shows both buying intent and relevance. Most intent is recognized before that point by a lexicon, the conversation structure and a small local classifier trained on model labels.

How it connects

  1. 01 Discover sources
  2. 02 Scan messages
  3. 03 Normalize & dedupe
  4. 04 Intent gate
  5. 05 LLM confirmation
  6. 06 Deliver card

What we delivered

  • Telegram client layer with an account pool, operation budgets and flood-wait handling
  • Source discovery from links, forwards and web catalogs, with a qualification funnel and a saturation metric
  • Normalization of look-alike and invisible characters, language detection and four-level deduplication with SimHash
  • Exact matching with an Aho–Corasick automaton and safe RE2 rules, plus semantic search with pgvector
  • AI cascade with a two-axis intent and domain gate, budget and pacing control, and cost tracking per provider
  • Brand-mention registry with sentiment
  • Delivery through a Telegram bot: explained cards, feedback buttons, one-time access codes, roles and an audit trail
  • More than 700 automated tests

Functionality

  • A local intent classifier, trained on model labels, keeps model calls to a small share of messages
  • Provider layer for Claude and Gemini with cost tracking
  • Benchmarked matching throughput
  • JSON Schema contracts between pipeline stages
  • A background service keeps the scanner running, raises an alert when it goes quiet and restarts it

Integrations

  • Telegram client API (MTProto)
  • Telegram Bot API
  • Anthropic Claude and Google Gemini APIs

Technologies

  • Python 3.12
  • Telethon
  • PostgreSQL + pgvector
  • asyncpg
  • Redis
  • lingua
  • pyahocorasick
  • google-re2
  • pytest

Result

Running as a pilot on real communities since August 2026 and delivering cards to the team. Account-pool capacity is the current limit; a multi-tenant version with an API is the next phase.

How do you keep model costs under control?

By making the model the last step, not the first. Keyword and semantic matching narrow the stream to messages that are on topic. A lexicon, the shape of the conversation and a small local classifier then decide whether a message shows buying intent. Only messages that pass both checks reach the language model, which confirms the intent and drafts a reply. Every model call is budgeted and logged with its cost, so spend is known before the month ends.

Discuss your project

Tell us about your project