Skip to content
Kenya 2027Discourse monitor

Measurements of a detection method and the pipeline behind it, ahead of the 2027 election.

This page shows how well a method works, not how much manipulation there is.

We collect Kenyan political discourse from X and run a method that looks for accounts acting together. What is defensible to publish about that is narrow: whether the method reproduces published results, how often our Kenya filter is wrong, what the detector surfaced and how a reader judged it, where it failed its own tests, and whether the pipeline is running. It is not a rate, and it names no account.

01 / Method

The method reproduces published results on 5 of 6 benchmarks

BasisMacro-F1 of the unsupervised IOHunter-style detector re-implemented here, on the 6 published IO benchmark datasets, against the number the paper reports. A claim about this copy of the method, never a Kenya result.

5 of 6 datasets land within 0.2 Macro-F1 of the number the paper reports. The exception is Iran, where the benchmark disagrees with itself.

Macro-F1 per benchmark dataset, paper against ours
DatasetPaperOursDifference
UAE84.6684.64-0.02
cuba57.9257.920.00
russia87.6587.83+0.18
venezuela95.0595.050.00
iran60.8371.43+10.60
china63.6663.67+0.01

iran: the benchmark authors' own code, run on their own release, gives 71.31 where the paper says 60.83. Ours is 71.43, which is +0.12 from the reference code. Iran is judged against the reference implementation's own output (71.31), not the published 60.83, because the benchmark disagrees with itself.

Source: docs/analysis/v2-findings.md / Measured 2026-09-11

02 / Filter

Our Kenya filter is wrong in measurable ways

Basis88 posts drawn fresh, labelled by a human blind to both methods, and corpus-weighted. Precision: of the posts the filter calls Kenyan, the share that are. Recall: of the Kenyan posts, the share the filter finds.

The learned classifier finds about 83% of Kenyan posts at about 89% precision. The keyword gate has precision 0.973 and recall 0.516, so it misses about 48% of Kenyan posts.

Relevance filter precision and recall on 88 labelled posts
MethodPrecisionRecall
keyword gate0.9730.516
learned classifier0.8890.827

Human labels, drawn fresh, judged blind, corpus-weighted. The honest framing is "the filter finds about 83% of Kenyan posts at about 89% precision", not "X% of the corpus is Kenyan".

Source: analysis/investigations/2026-09-12-relevance-classifier/findings.md / Measured 2026-09-13

03 / Output

What the detector surfaced, and what a reader made of it

BasisOne frozen snapshot, not the live run. Each method's communities were read blind by a model reader at a relevance floor of 0.5; cases are communities read, and the counts are how many that reader called political or Kenya-relevant. This describes what the method surfaced. It is not a rate: there is no denominator.

Communities read blind, by method
MethodCommunities readPoliticalKenya-relevant
v217512
v11713

Neither method surfaced anything the reader called an influence operation. Snapshot 2026-09-05-promotion-off__emb20260912.

Model verdicts, 17 cases per method, read blind. A description of the method's output, never a prevalence rate - there is no denominator. The latest daily run (10 Oct 2026, 21:20 EAT) produced a listing. The listing is account level, so it is not published and its length is not shown.

Source: analysis/investigations/2026-09-12-component-ranking/findings.md / Measured 2026-09-12

04 / Failures

Where the method failed its own tests

BasisEach is a measurement of this project's own method failing, so no collection artifact can flatter it.

  1. A high self-amplification score does not separate engagement pods from influence operations. Confirmed operations self-amplify more than the pods this project's detector found, so the filter was dropped.

    effect size d = +0.78, confirmed operations above pods

    Standardised mean difference of the self-amplification score, confirmed operations against the pods found on this corpus. Group sizes are not recorded in the cited document.

    Source: docs/OBJECTIVES.md (A3)

  2. At a cosine cut of 0.85 in this encoder, most admitted post pairs are unrelated Sheng replies, not copies.

    59 of 100 sampled pairs unrelated; 26 same message; 15 same topic

    100 cross-author post pairs at or above 0.85 drawn from 279,060 behind the top 500 and labelled blind by a reader. Of all 279,060 pairs, 3% were near-copies and the median pair shared no words.

    Source: docs/analysis/v2-findings.md (section 4) / Checked 2026-09-11

  3. A raw toxicity series over this corpus is not publishable: it moved because the collector changed what it collected, not because the discourse changed.

    raw series +78%; composition-standardised series flat then falling

    Baseline scope, July to August 2026. The baseline mix moved from 70.7% search and 12.1% replies to 14.3% search and 72.6% replies. The standardised series holds the mix at the earlier reference week.

    Source: docs/analysis/2026-09-13-publishable-statistics.md

Source: docs/analysis/2026-09-13-publishable-statistics.md / Measured 2026-09-13

05 / Pipeline

Is the pipeline running?

Daily detector run

BasisThe newest persisted run of the daily detector, from its own run record. Stale means older than 36 h when this page was built. The listing's length and contents are not published.

Last run
10 Oct 2026, 21:20 EAT
Window
21 days to 2026-10-09
Run id
20261010T182001Z
Runs recorded
2
Run days, last 14
2026-10-10
Missed days
none
networks (s)
325.0
fuse centrality (s)
2.9
relevance (s)
61.4
communities (s)
5.3

Source: coord2/kind=daily_runs

Collection volume

BasisRows the collector wrote to posts/ per UTC collection day (the dt partition), all post types. Includes re-collections of the same post, so it is not distinct posts and not Kenyan posts. Scope is the partition's own type, not first-seen type. A day with no partition is a gap day: nothing was collected and it cannot be backfilled. Today's open partition is excluded.

Collection began on 2026-07-16 (UTC). Across the whole period 13 days have no data at all. A missed day is a hole: X search reaches back 14 days, so it cannot be collected after that.

0125,000250,0002026-09-10: 68,567 rows2026-09-11: 51,903 rows2026-09-12: 50,209 rows2026-09-13: 84,270 rows2026-09-14: 206,119 rows2026-09-15: 105,457 rows2026-09-16: 128,658 rows2026-09-17: 107,704 rows2026-09-18: 103,788 rows2026-09-19: 122,511 rows2026-09-20: 100,140 rows2026-09-21: 73,355 rows2026-09-22: 66,848 rows2026-09-23: 67,694 rows2026-09-24: 62,806 rows2026-09-25: 91,716 rows2026-09-26: 63,514 rows2026-09-27: 63,092 rows2026-09-28: 95,337 rows2026-09-29: 56,538 rows2026-09-30: 72,967 rows2026-10-01: 98,611 rows2026-10-02: 72,387 rows2026-10-03: 63,442 rows2026-10-04: 80,943 rows2026-10-05: 58,107 rows2026-10-06: 57,506 rows2026-10-07: 82,338 rows2026-10-08: 63,425 rows2026-10-09: 50,031 rows09-1009-2510-09
  • baseline
  • targeted
  • control
  • other
  • gap day, nothing collected
Show the numbers
Rows written per UTC collection day by scope
UTC dayRowsObjectsbaselinetargetedcontrolother
2026-10-0950,0317942,5226,1861,3230
2026-10-0863,4258854,6337,7161,0760
2026-10-0782,33812974,4667,0028700
2026-10-0657,5069748,2868,2609600
2026-10-0558,1079051,0315,8791,1970
2026-10-0480,94312972,1337,2241,5860
2026-10-0363,4428355,6086,6131,2210
2026-10-0272,3879364,2926,6991,3960
2026-10-0198,61113390,5906,3221,6990
2026-09-3072,9678865,5336,0591,3750
2026-09-2956,5389747,5927,1861,7600
2026-09-2895,33714284,3419,2921,7040
2026-09-2763,09210553,9517,8991,2420
2026-09-2663,51410653,5608,7381,2160
2026-09-2591,71613082,6727,8111,2330
2026-09-2462,80610252,4409,0651,3010
2026-09-2367,6949858,7657,5181,4110
2026-09-2266,8488963,0903,5312270
2026-09-2173,35510265,0287,6736540
2026-09-20100,14014068,63830,9475550
2026-09-19122,51117098,62323,1996890
2026-09-18103,78814366,42736,4658960
2026-09-17107,70413870,45936,2331,0120
2026-09-16128,658176104,17223,6118750
2026-09-15105,45712769,02536,0114210
2026-09-14206,11922970,746134,9734000
2026-09-1384,27012578,5495,72100
2026-09-1250,2096843,6646,54500
2026-09-1151,9036744,1067,79700
2026-09-1068,5679458,6159,95200

Source: posts/ partitions: object listing and parquet footers

Enrichment lag

BasisLatest dt partition per stage, and rows written to it over the last 7 closed UTC days. A stage's dt is when it ran, not the date of the posts it scored, so rows are throughput and not coverage. lag_days is the newest posts dt minus the stage's newest dt. relevance is the promoted model only.

Newest partition, lag and recent rows per enrichment stage
StageNewest dtLag (days)Rows, last 7 d
posts2026-10-100455,792
embeddings2026-10-10025,500
relevance2026-10-100220,160
hatespeech2026-10-10025,500
incitement2026-09-28120

Source: posts/, embeddings/, relevance/, hatespeech/, incitement/: listings and footers

07 / Not shown

What this page leaves out

Each of these is a number or a claim this project could produce and has decided it cannot defend. They are listed so the gap is visible, not so a reader fills it with an assumption.

  • Any prevalence rate

    Nothing in the collector samples the discourse at random, so no rate over it has a denominator. A control arm is needed first.

  • Any account, handle or per-account claim

    Detection says accounts act together; neither it nor a reader's judgement establishes that a named account is inauthentic.

  • Counts of coordinated accounts or communities from the live run

    The listing length is a triage budget and the community count depends on a clustering resolution and seed. Neither is a finding.

  • A toxicity series over time

    The raw series moved with the collector's own target list. Only a composition-standardised or within-partition rate may be shown, and none is shown yet.

  • Cross-channel corroboration of clusters

    The earlier statement about it is false and was retired.