B2B SaaS · channel & agency partnerships

Agency partner sourcing engine: resumable harvesting of B2B directories for a partner program

A sourcing engine that goes deeper into public B2B agency directories with every daily run, resolves each agency's own website, verifies contacts and feeds an isolated partner-recruitment sequence.

Client
Client confidential
Industry
B2B SaaS · channel & agency partnerships
Year
2026

Challenge

The same SaaS client needed marketing agencies and solution providers to join its partner program. Public B2B agency directories are large and deeply paginated, and a crawler that starts from the first pages every day keeps finding the same agencies. Directory profiles also mix agency websites with tracker, CDN and social links. The partner outreach had to run fully apart from the client's startup outreach, with its own sender reputation.

Solution

We reused the outreach core of the startup platform and added a sourcing layer. It harvests agency profiles from public B2B directories, resolves each agency's own website, then crawls it, extracts contacts and verifies them. Competitors and free-mail domains are excluded automatically. Qualified contacts enter an A/B-tested first-touch sequence with up to three follow-ups. The instance runs on its own database, dashboard, sending domain and sender identity.

How it connects

  1. 01 Harvest directory profiles
  2. 02 Resolve agency websites
  3. 03 Crawl & extract contacts
  4. 04 Filter & verify
  5. 05 Approval buffer
  6. 06 First touch + follow-ups
  7. 07 Reply tracking

What we delivered

  • Directory connectors with persistent per-source crawl cursors
  • Rotation across country and category combinations, so each daily run goes deeper instead of rereading the first pages
  • Website resolution, contact extraction and verification
  • Automatic exclusion of competitors and free-mail domains
  • A/B-tested first-touch sequence with up to three follow-ups
  • Isolated deployment with a separate database, dashboard, sending domain and sender

Functionality

  • Noise filtering that keeps tracker, CDN and social-network domains out of the lead base
  • Domain normalization that collapses subdomain variants into one company
  • Transient-failure handling: a timeout never resets the crawl position
  • Warm-up limits, a same-day approval buffer and random delays between sends
  • Reply tracking, opt-out handling and role model inherited from the core platform
  • Global suppression list shared with the client's other outreach instances

Integrations

  • Public B2B agency directories
  • Headless-browser rendering for directory pages
  • Email-verification API
  • SMTP sending and IMAP reply tracking

Technologies

  • Python 3.12
  • FastAPI
  • SQLAlchemy 2 + Alembic
  • PostgreSQL 16
  • Jinja2
  • httpx + BeautifulSoup
  • APScheduler
  • Playwright (headless Chromium)
  • Docker Compose
  • Nginx + Let's Encrypt

Result

In production, with the sourcing layer extended since launch. Crawl cursors persist per source, so each daily run continues where the previous one stopped, and the partner program's sender reputation stays separate from the client's other outreach.

The engine runs on the same outreach core as the startup discovery & outreach platform and adds its own sourcing layer on top.

Discuss your project

Tell us about your project