A deep dive into kaito.ai (Mindshare / Yaps / Voices / Pro) and an honest feasibility case for building an EyezOn-owned attention/mindshare layer on the stack we already have — zero Kaito dependency.
Kaito turned one idea into a ~$33M/yr business: attention is a leading indicator, so measure it as a number. They ingest crypto social at scale, LLM-score every post for information value, and publish a mindshare metric (share-of-voice, influence-weighted) plus a Voices leaderboard of the most influential KOLs per vertical (trading / AI / crypto).
The EyezOn angle: we already own the two things that make this expensive for everyone else — a working free X reader (the @day1meme burner, no paid API) and an accumulating proprietary social + on-chain outcome dataset. That means we can build a narrower but deeper mindshare layer for the one vertical that pays our bills (Solana trading / memecoins), and validate it against something Kaito structurally lacks: our own early-call accuracy data. Kaito can tell you who is loud. We can tell you who is loud and early and right.
Feasibility: realistic — as a scoped internal signal layer first, product surface second. Not a 1:1 Kaito clone (their 200M-posts/day firehose is out of reach on a burner), but the useful 20% — mindshare-per-token/narrative + a validated trading-Voices leaderboard — is buildable on our current stack in phases, at near-zero marginal cost, replacing a $750/mo+ rented dependency.
Kept near the top on purpose — this is the "so what."
Kaito (kaito.ai, token KAITO) positions itself as "the data layer for the attention economy" / "the Bloomberg of crypto." It is fundamentally an InfoFi (information-finance) company: measure crypto attention, then monetise the measurement. Products:
The core metric and product. Mindshare = a token / project / narrative's share of total crypto attention over a timeframe, expressed as a percentage and in basis points (1 bps = 0.01%). Not raw mention count — it's share-of-voice among top-ranked posters, weighted by content quality, relevance/focus, originality and the influence of who engaged. Tracked over time on leaderboards so you can watch attention rotate between narratives.
Launched Dec 2024. Daily points ("Yaps") rewarded to creators ("yappers") for high-quality crypto posts on X. Projects (Berachain, Story, dYdX, Injective, Mitosis…) ran project-specific yapper leaderboards to allocate airdrops. The scoring bet: AI can score the information value of a post — did it surface real alpha, did the right people engage — rather than vanity follower/retweet counts. Note: this product was hit hard by X's Jan-2026 API crackdown on pay-to-post apps (see §6).
A leaderboard ranking the most influential individual voices whose activity moves discourse, segmented by vertical (the linked page: Trading, alongside AI and Crypto). Columns observed live: Mindshare (that voice's share of attention in the sector), Δ Mindshare (change, in bps and %), Smart Followers (SF) count, and Δ Smart Followers. Filterable by timeframe (7D / 30D / 3M / 6M / 12M) and by "Top Gainer." This is Kaito's answer to "who are the KOLs that matter in this niche, right now, and who's rising."
The paid terminal / "Bloomberg of crypto." MetaSearch across thousands of Web3 sources by ticker/topic/trend; Token Mindshare + Narrative Mindshare tracking; sentiment analytics; smart alerts & custom feeds; a Catalyst Calendar (unlocks, TGE, governance, tokenomics events); an audio library of podcast/conference transcripts with AI summaries; and an AI Copilot. Institutional API access on top. This is the ~$33M/yr revenue engine.
After the Yaps crackdown, Kaito split into Kaito Studio (tier-based, brand-selected creator marketing across X/YouTube/TikTok/Instagram) and Kaito Markets — a Polymarket partnership turning mindshare scores into prediction-market assets ("Will Anthropic's mindshare exceed OpenAI's?"). One pilot market drew >$1.3M in wagers. Strategically this reduced their single-platform dependency by pulling from many platforms at once — a lesson for us.
Kaito keeps the exact algorithm deliberately opaque (to resist farming), but between their docs, the MCP server schema, and coverage, the architecture is clear. It's a fairly standard social-listening + LLM-scoring + aggregation stack, just executed well and at scale.
KAITO PIPELINE (reconstructed)
1 · INGEST 2 · CLASSIFY (AI/LLM) 3 · WEIGHT 4 · AGGREGATE 5 · SERVE
X / Twitter ─┐ ┌─ token / project tag ─┐ author influence ─┐ share-of-voice ─┐ Mindshare leaderbds
Telegram ─┤ ├─ narrative / vertical ─┤ (smart-follower ├ per token / ─┤ Voices / Arena
Farcaster ─┼──▶ ├─ sentiment (bull/bear) ─┼─▶ social graph) ┼─▶ narrative / ─┼─▶ Yaps points
Governance/ ─┤ ├─ quality / insight ─┤ engagement qual. ─┤ vertical, per ─┤ Pro terminal + API
forums/Medium ─┤ ├─ originality (plagiar.) ─┤ (WHO engaged) ─┤ timeframe ─┤ sentiment series
podcasts/news ─┘ └─ spam / bot filter ─┘ time-decay ─┘ dedup + norm ─┘ Attention Markets
~200M crypto posts/day the "smart follower" graph is the secret sauce
At peak, ~200M crypto-adjacent posts/day across X, Telegram and Farcaster, plus "tens of thousands of Web3 sources": governance proposals/forums, Medium, conference & podcast transcripts, news, and research reports. Heavy dependence on the X API for the firehose — which became their single point of failure.
Every post is machine-tagged and scored. Inferred sub-tasks, each of which we'd replicate:
Raw tagged posts are weighted before aggregation. The dominant factor is the Smart Follower graph: an AI-identified network of reputable crypto accounts. Engagement from a smart follower is worth far more than the same engagement from a low-influence/bot account — so who engages beats how many. Additional weights: relevance/focus (a post about the token > a market roundup that mentions it), engagement quality, originality, posting consistency, and — critically — a claimed downstream-behavioral signal (did readers actually go trade). Everything is time-decayed, so mindshare reflects current attention.
Weighted, deduped, bot-filtered post signals are summed into a share-of-total-impressions for each token / narrative / vertical, per timeframe, then normalized to a % / bps figure. "Voices" ranks individual authors within a vertical by their own influence + mindshare contribution (and shows their smart-follower count and its delta).
Leaderboards (Mindshare, Arena, Voices), Yaps points, the Pro terminal, sentiment time-series, and the API/MCP surface (§4). The algorithm is intentionally undisclosed — a tell that the weighting recipe, not the raw data, is the moat.
The public Kaito MCP server (MetaSearch-IO/kaito-mcp-server) exposes their product cleanly. This is the useful-20% target list for our own build — replicate the parts that matter for trading, skip the rest:
| Kaito capability (tool) | What it returns | Do we need it? |
|---|---|---|
kaito_mindshare_entity | Daily mindshare time-series for a token | core — build |
kaito_mindshare_entity_arena | Projects ranked by mindshare score | core — build |
kaito_mindshare_entity_delta | Top gainers/losers by mindshare change | core — build |
kaito_mindshare_narrative | Daily narrative-level mindshare series | core (= meta-flow) |
kaito_mindshare_entity_by_account | Top KOLs ranked by a token's mindshare | core (= Voices) |
kaito_smart_followers / _following | Smart-follower counts/gains; who a user follows | approximate |
kaito_sentiment_entity | Daily bull/bear volume-weighted sentiment + events | build (LLM) |
kaito_tweet_engagement_info | Likes/RT/replies/views + "smart engagement" count | reader gives this |
kaito_engagement / kaito_mentions | Daily engagement + mention counts per token/keyword | build |
kaito_search / kaito_advanced_search | NL + structured search across Twitter/News/Research/Podcast | partial (X only) |
kaito_feeds | Top-ranked content feed, market or per-token | = Attention Feed |
kaito_entities / kaito_narratives | Token & narrative resolvers (taxonomy) | build small |
kaito_events | Catalyst calendar (unlocks/TGE/governance) | skip for now |
kaito_twitter_user_metadata | Profile/follower stats + account classification | reader gives this |
The centrepiece. We are not cloning Kaito's firehose; we're building a scoped, validated, owned mindshare layer for Solana trading, on primitives we already have. The unfair advantage is that steps 1 and 6 already exist in some form.
EYEZON OWNED MINDSHARE LAYER ① INGEST (owned) ② CLASSIFY (LLM, cost-cut) ③ WEIGHT + SCORE x_reader.py @day1meme burner per post → author influence score • curated KOL/voice list • token/CA tag (resolve vs our CAs) (followers × smart-eng × • SearchTimeline $cashtags • narrative / vertical tag our history of early hits) • per-scan on-demand reads • sentiment (bull/bear) engagement quality (WHO) fallbacks: t.me/s, fxtwitter • quality / insight threshold time-decay (recent = heavier) + source-virality seed post • spam / caller-bot filter bot/wash filter (vol-bot adv.) │ │ Haiku-route + prompt-cache + Batch │ ▼ ▼ ▼ ④ AGGREGATE (our DB, bounded) ⑤ RANK VOICES (trading) ⑥ VALIDATE (our moat) share-of-voice per: influence × engagement × join each voice's early • token / CA relevance × recency ───────────▶ mentions to OUR on-chain • narrative (meta-flow) = Voices leaderboard outcomes (calls_by_ca, • vertical Perception perf DB) normalized %/bps, time-decayed → CALL-ACCURACY weight retention caps on every table (Kaito can't do this) │ │ │ └────────────────────────┬───────────────┴────────────────────────────┘ ▼ ⑦ SURFACE (a) feeds Perception attention/virality + KOL-cluster + 🐳+👁 (b) product: "EyezOn Mindshare" + "Top Trading Voices" leaderboard
x_reader.py + @day1meme burner cookie reads any account's feed & engagement free — no paid X API. This is the ingest layer Kaito pays dearly for.calls_by_ca, Perception perf DB (265 real calls), whale DB, our own candle engine — the ground truth for call-accuracy validation.Point the existing reader at ~50–150 curated trading voices, pull recent posts for a handful of live tokens, LLM-tag token+sentiment on a small batch, and eyeball whether a crude share-of-voice number tracks moves we already know. Goal: confirm signal exists before any storage/scale investment. Pure notebook work, throwaway.
Formalize the LLM tagger (token/CA, narrative, vertical, sentiment, quality, spam) with the cost-cut playbook baked in from day one. Compute per-token share-of-voice (time-decayed, bot-filtered) into a bounded table. Wire it as a new input to Perception's attention/virality components (shadow first, don't touch live scoring). Retention caps ship with the table.
Aggregate to narrative level = the Zenchat meta-flow detector: which meta is forming (narrative clustering + KOL attention velocity) before it's named, and attention deceleration to call "the 3rd runner is the last." This is the piece Kaito arguably does worse than we can, because we fuse it with fresh-token on-chain clustering.
Rank voices per vertical by influence × engagement × relevance × recency — then the differentiator: join each voice's early mentions to our on-chain outcome data and weight up voices who were repeatedly early on real winners. Output = a Voices board that ranks not by who's loudest but by who's been right first. Feeds the KOL-cluster trigger and 🐳+👁 fusion directly.
Expose "EyezOn Mindshare" + "Top Trading Voices" in the social-scanning layer (Radar Attention Feed, N-eyes, rank ladder). Free = counts; Pro = full breakdown + validated voice ranks. Natural paywall, on-brand with the existing roadmap. Product only after the internal signal proves out.
Not sugar-coated — these are the real constraints and the realistic mitigations/bridges.
Kaito ran 200M posts/day via the paid API. Our burner cannot and must not try to — X fingerprints read cadence even read-only, and a banned burner kills ingest. Mitigations: (a) scope hard to Solana trading, not all of crypto — we need hundreds of curated accounts, not the firehose; (b) low-and-slow, jittered read cadence to protect the burner (already the standing rule); (c) prioritise reads — a curated voice list + on-scan reads + $cashtag search, not blanket timeline crawling; (d) fallbacks already known (t.me/s mirrors, fxtwitter) if the cookie dies; (e) a small pool of burners as a bridge if one list is insufficient (ask Pete before adding — anti-ban hygiene). Honest ceiling: this gives us depth on a curated set, not exhaustive coverage. That's the right trade for a trading-vertical edge, but we should never claim "total" mindshare — ours is "mindshare among the voices that matter," which is arguably more useful anyway.
Tagging every post with an LLM is the cost center. Left naive it would blow the budget. Mitigations (the standing cost-cut playbook, non-negotiable in step ②): prompt-cache the standing tagging instructions; Haiku-route the trivial/high-volume calls (most posts are easy classifies) and reserve bigger models for genuine ambiguity; Batch API for the non-real-time daily aggregation passes → 60–90% off with no quality hit. Cheap pre-filters (regex/keyword/cashtag + engagement floor) cull the obvious noise before any token is spent. Groq/Cerebras open models are an option for the narrow, high-volume tagging subtask only — not a Claude replacement.
Every data-writing part ships with retention/caps — this is a hard standing rule after the 2026-07-21 Railway volume-fill incident (a signatures table hit 9M unbounded rows, volume 86% full). For this layer: cap raw-post retention (aggregate then prune, e.g. keep computed daily mindshare + a short raw window), TTL on the KOL→token map's cold entries, and storage monitoring so I catch a fill early, never Pete from a dashboard. DuckDB/Parquet for the aggregate time-series tier; SQLite single-writer+WAL if it lands near eyezon_data.db.
The exact failure that killed Yaps: once a score is visible and rewarded, people farm it. Two forms hit us — (a) caller-bot/gem-alert spam manufacturing fake $cashtag volume after a pump (proven on $HNUT: 37 bot posts polluted the naive scan), and (b) sybil/wash if we ever expose it as a public product. Mitigations: weight by author reputation + smart-engagement not raw counts (same insight as Kaito's smart-follower graph); source-virality traces the meme to its original high-reach seed post rather than trusting downstream cashtag volume; reuse the vol-bot adversary filter; and if it becomes a product, reputation-weighting + cost-to-signal + anomaly detection (already flagged as the risk in the social-layer roadmap). Keep our scoring recipe private — like Kaito, the weighting is the moat.
X actively hunts InfoFi/attention apps. Our profile is quieter (read-only, curated, low cadence, not paying users to post) so we're not the Yaps target — but not zero risk. Mitigation: never use Pete's main account; treat the burner as disposable; keep fallbacks warm; and, longer term, diversify sources (Telegram mirrors, Farcaster) exactly as Kaito was forced to after the crackdown — better to design multi-source early than get single-platform-killed.
*Enterprise single-seat, billed biennially per Pete's screenshot; MCP access included on all plans.
Marginal cost dominated by LLM tagging; near-zero at our scoped volume with Haiku-routing + Batch.
The ROI isn't only the ~$9k/yr saved. It's that the rented version can't feed our engine or our product (no owned data, no API integration without enterprise contact-sales, no call-accuracy validation), whereas the owned version compounds: every day it runs, our KOL→token map + call-accuracy weights get better and less reconstructable by anyone else. Renting is a cost; owning is an appreciating asset. Textbook north-star.
Feasible. Build the scoped internal signal layer, not the Kaito clone.
Per Pete's scope: this is research + feasibility only. No build, prototype, or deploy has been done — this report is the deliverable.
EyezOn internal research · read-only · prepared 2026-07-22 · noindex. Kaito figures (posts/day, revenue, pricing, dates) are from third-party coverage and Pete's pricing screenshot; treat as directional, not audited. This document proposes nothing beyond feasibility — no build authorized.