Autonomous Competitor & Market Intelligence
AI Workflow Pipeline · n8n, OpenAI & Deterministic Scoring Engine
An end-to-end competitive intelligence pipeline orchestrating scheduled website polling, 0-token content hashing, GPT-4o semantic fact extraction, deterministic mathematical priority scoring, and dual-track delivery to Telegram and executive dashboards.




Pipeline Overview & Philosophy
The flow is straightforward: competitor websites, pricing pages, and changelogs are polled on schedule, n8n orchestrates the pipeline, OpenAI handles semantic reasoning, and the output feeds real-time Telegram alerts and a weekly executive briefing dashboard.
The core engineering challenge lies in deciding where different types of processing belong: what should be strictly deterministic, where an LLM adds actual value, and how to avoid notification fatigue for leadership. Working at the intersection of design and data systems, the most effective AI applications are usually around 80% deterministic plumbing and 20% model reasoning. The real value is orchestrating the flow so the model only touches what code cannot solve on its own.
Core Architectural Pillars
1. The 0-Token Gate
Rather than sending every crawled webpage straight to an LLM, the workflow cleans and hashes the content first. If the computed checksum matches the previous snapshot, the execution halts immediately. Over 85% of runs terminate right here—meaning zero LLM token waste on pages that have not changed.
2. Separating Reasoning from Scoring
LLMs are strong at semantic extraction (identifying pricing updates, product pivots, or regional expansion), but should not decide organizational priority on their own. The system separates the two stages: GPT-4o extracts structured facts, while a deterministic scoring engine applies mathematical weights based on event category, competitor tier, and magnitude before deciding if an alert is warranted.
3. Real-Time Alerts vs. Macro Trends
Individual competitive moves and broad market shifts serve different purposes, so they run in separate pipelines. Urgent commercial changes (like a 40% enterprise price hike) trigger instant Telegram alerts for sales and GTM teams, while a weekly pipeline analyzes a trailing 30-day window to synthesize macro strategic patterns.
4. Enforcing Anti-Hallucination Rules
LLMs have no problem inventing connections without factual grounding. Hard constraints are imposed to avoid noise. The model is required to cite exact before/after text snippets, and for weekly trends it is programmatically disallowed to suggest any pattern that does not refer to at least two distinct, verified events.
Pipeline Stages & Separation of Concerns
- Ingestion & Polling: Scheduled scrapers fetch competitor websites, pricing pages, and changelog feeds on continuous intervals.
- Hash Gating: Computes SHA checksums over normalized DOM content, short-circuiting over 85% of executions when no diffs exist.
- Semantic Extraction: Constrained GPT-4o prompt schema extracts strictly verifiable before/after changes with direct citations.
- Deterministic Scoring: Mathematical evaluation engine weighs competitor tiers, category impact, and magnitude to prevent notification fatigue.
- Presentation & Routing: Instant Telegram webhook notifications for commercial shifts and weekly executive briefing summaries.
Outcome & Evolution
The system continues to evolve, but the strict separation of ingestion, hash gating, semantic reasoning, and presentation delivers reliable intelligence while keeping operating token costs minimal and maintaining high signal-to-noise ratios for leadership.
Contact me
Have a question or want to work together?
Drop me a message.