{"i":9,"slug":"system-one-arena","chart":{"id":"arena-accuracy","title":"Who decides right? Accuracy on 1,000+ checkable decisions","subtitle":"Share of graded items answered correctly, first presentation · 95% Wilson intervals","kind":"dot-range","unit":"rate","whisker":"ci95","series":[{"name":"Accuracy","points":{"$k":["label","value","lo","hi","n","highlight"],"$r":[["Jev 1.13",0.7681,0.7416,0.7927,1048,true],["Clef 27B",0.6994,0.671,0.7264,1048,"\u0001"],["Clef-Flash 9B",0.6517,0.6224,0.68,1048,"\u0001"],["Kev 4B",0.6403,0.6107,0.6688,1048,"\u0001"],["lev 4B",0.5859,0.5558,0.6153,1048,"\u0001"],["Laya",0.2739,0.2477,0.3016,1048,"\u0001"],["Julia-1",0.2586,0.233,0.2859,1048,"\u0001"]]}}],"note":"1048 graded items in five suites. An error, a timeout or a label outside the option set counts as wrong. Two models differ clearly only where the paired McNemar test says so (table below).","sourceIds":["system-one-arena"]}}