← Back to Tech Radar
Hacker News tech

Can a MUD evaluate LLMs? A $99 proof of concept

Trending on Hacker News: Can a MUD evaluate LLMs? A $99 proof of concept (0 points, via cruciblebench.ai)

In one line

CrucibleBench places language models in a persistent text world, a MUD where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 turns.

Opening excerpt

Nintendo’s Gunpei Yokoi used the phrase to describe a design philosophy: take mature, inexpensive, well-understood technology and use it in a new way. CrucibleBench applies it to AI evaluation.

We did not choose a MUD because it is charming. We chose it because its constraints make behavior measurable.

Static benchmarks measure what models know in isolation. They do not measure how models behave where trust must be earned, information is gated by relationships, and blunt questioning raises suspicion.

7 command types, 12 rooms, 14 items. Hallucinated actions and wrong-room interactions are detectable, and action efficiency is measurable.

(Excerpted from the original; full article via the source link below.)

Source: Hacker News

Related Services

Related Reading