Early standings
How people rate world models after exploring them in real time. Rankings account for both approval and the number of votes.
7 votes recorded across the models belowRates show the share of “Yes” votes. Rank also accounts for the number of votes. How rankings work
What these ratings measure
A world model generates a scene that changes as you move and look around. A session lets you experience its visual quality, how consistently the scene holds together, and how it responds to your controls.
After each session, the arena asks one question: “Did you like this model?” The model’s identity is revealed after you answer, to reduce the influence of its name or maker. Each vote is an overall impression of one model in one session; the individual qualities above are not scored separately.
Reading the table
Rate is the percentage of recorded votes that were “Yes.” Runs counts sessions with a vote, rather than every session started. A dash means the model has no votes yet. Skipped votes and sessions that never produced a world do not contribute to its score.
The approval bar shows uncertainty from the number of votes. A longer faded segment means a less precise estimate. As votes accumulate, that range generally narrows.
How models are ranked
A model with one positive vote has 100% approval, but very little evidence behind it. To account for sample size, we rank models by the lower end of their 95% Wilson confidence interval: a conservative estimate of approval based on the votes received.
This is why a model with four positive votes out of five ranks above one with a single positive vote, even though its displayed rate is lower. The solid part of the bar ends at the value used for ranking; the faded segment extends to the interval’s upper end. The small marker shows the observed approval rate.
Interpreting the results
Visitors choose their own scenes and can rate multiple sessions. The totals therefore represent votes, not unique participants, and each model may be evaluated on a different mix of scenes and session lengths.
These standings describe approval among the sessions rated here. Small differences deserve caution, especially when vote counts are low. The confidence bars account for sample size; they do not account for differences in who participated or which scenes they chose. The results will evolve as more sessions are rated.