Live Platform Data
Drux Activity
Real-world AI model performance from every query run on Drux. Speed, reliability, and consensus measured on actual user questions — not controlled benchmarks.
998 searches · 3,045 model calls · updated every 5 min
17.1Mtokens processed in the last 30 days
Platform Summary
998 total searches completed. 3045 individual model calls recorded. Average consensus score: 63%. Model completion rate: 74%.
Top Models by Speed
- inception/mercury-2: avg 4.6s, 97% reliability, 194 queries
- google/gemini-2.5-flash: avg 7.0s, 40% reliability, 5 queries
- mistralai/mistral-small-2603: avg 7.9s, 99% reliability, 199 queries
- microsoft/phi-4: avg 8.5s, 99% reliability, 217 queries
- x-ai/grok-4.3: avg 9.8s, 100% reliability, 12 queries
Most Agreed Questions
- What are the 5 C's of effective negotiation? — consensus score 100%
- What’s healthier, white rice or brown rice? — consensus score 90%
- What is healthier, brown rice or white rice? — consensus score 90%
- Is classical AI research still being done because it helps modern AI? — consensus score 90%
- Is it true goldfish have a memory span of only three seconds? — consensus score 90%
Platform Pulse
Total Searches
998
This Week
96
last 7 days
Today
36
Model Calls
3,045
across all searches
Avg Consensus Score
63%
across all completed searches
Model Completion Rate
74%
calls that returned successfully
Model Leaderboard — on Drux
Last 998 searches · click column to sort
| Model | Avg Speed▲ |
|---|---|
1Mercury 2FREE | 4.6s |
2Gemini 2.5 FlashFREE | 7.0s |
3Mistral Small 4FREE | 7.9s |
4Phi-4FREE | 8.5s |
5Grok 4.3PAID | 9.8s |
6Command RFREE | 10.1s |
7Llama 4 MaverickFREE | 13.0s |
8GPT OSS 120BFREE | 13.2s |
9Claude Sonnet 4.6PAID | 13.5s |
10Gemini 3.5 FlashPAID | 14.1s |
11Hermes 3 70BFREE | 14.6s |
12Solar Pro 3FREE | 15.9s |
13Gemma 4 31BFREE | 17.1s |
14DeepSeek V3.2FREE | 19.5s |
15Sonar ProPAID | 22.7s |
16GLM 5.2FREE | 23.9s |
17GPT-5.5PAID | 24.1s |
18Nemotron 49BFREE | 24.8s |
19GPT-5 MiniFREE | 25.3s |
20ERNIE 4.5 300BFREE | 27.8s |
21Seed 1.6FREE | 28.5s |
22Claude Opus 4.7PAID | 31.1s |
23Qwen3.5 397BFREE | 35.9s |
24OLMo 3 32BFREE | — |
⚡ Fastest on Drux
Mercury 2
avg 4.6s · 97% reliable
✓ Most Reliable on Drux
Gemini 3.5 Flash
100% completion · 38 calls
Consensus Digest — Recent Public Searches
✓ Models Agreed
What are the 5 C's of effective negotiation?
What’s healthier, white rice or brown rice?
What is healthier, brown rice or white rice?
Is classical AI research still being done because it helps modern AI?
Is it true goldfish have a memory span of only three seconds?
⚡ Models Diverged
I made a secure way for agents to request secrets from you using HyperDHT
Anyone got DeepSeek-v4-flash-0731 running using antirez/ds4?
What is the best approach for beginners to start learning?
Is ChatGPT or Claude better for general use?
Four Competing AIs Answer 1 Classic Question
Updated every 5 min · Data from last 1,000 searches
About This Data
How is this different from other AI benchmarks?
Most AI benchmarks are run under controlled laboratory conditions on standardised test sets. Drux Activity measures performance on real user questions — diverse, unpredictable, and representative of actual use. Speed and reliability numbers here reflect what users actually experience.
What is a Consensus Win?
When multiple models answer the same question and their responses agree closely (consensus score ≥ 7/10), all models that responded are credited with a Consensus Win. A high Consensus Win rate means a model consistently lands on the same answer as its peers — a signal of reliability beyond just accuracy.
Why do free models have more queries?
Drux randomly selects models for each search. Free-tier models are included in more searches because they are available to all users. Paid models are selected when users opt into premium tiers. Query count reflects availability and tier distribution, not quality — sort by Speed or Reliability for a fairer comparison.
How often is this updated?
The leaderboard updates every 5 minutes, pulling from the last 1,000 completed searches. As query volume grows the data becomes more statistically significant.