Models / Trinity Large Thinking

Arcee Ai

Trinity Large Thinking

13 published results from 2 sources. Each card shows where the number comes from and what it does not measure. Results are never combined into one score.

Provider
Arcee Ai
Sources
2
Our benchmarks
0
Price
Not yet published

Reported by others

1,371
Overall · rank 157 of 177
Unit
Arena rating, higher is better
Range
1,367 to 1,375
Sample
30120 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not an error rate: claims that cannot be checked on the web are skipped, and preference still carries most of the weight.
1,367
Business, management and finance · rank 170 of 402
Unit
Arena rating, higher is better
Range
1,359 to 1,375
Sample
6115 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,332
Creative writing · rank 168 of 407
Unit
Arena rating, higher is better
Range
1,323 to 1,342
Sample
4717 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,401
Expert prompts · rank 141 of 359
Unit
Arena rating, higher is better
Range
1,389 to 1,413
Sample
2903 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,357
Instruction following · rank 178 of 409
Unit
Arena rating, higher is better
Range
1,351 to 1,364
Sample
10346 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,368
Overall · rank 184 of 409
Unit
Arena rating, higher is better
Range
1,363 to 1,372
Sample
30277 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
1,343
Writing, literature and language · rank 176 of 408
Unit
Arena rating, higher is better
Range
1,336 to 1,351
Sample
7248 votes
Configuration
Trinity Large Thinking
Measured
25 Sep 2026
Not shown
Not accuracy or correctness: it ranks which answer voters preferred, with answer length and formatting controlled for.
Reported by OpenHands Index
32.1%
Average · rank 32 of 34
Unit
% resolved, higher is better
Configuration
Trinity Large Thinking (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
25.0%
Front end (SWE-Bench Multimodal) · rank 27 of 34
Unit
% resolved, higher is better
Configuration
Trinity Large Thinking (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
12.5%
Greenfield (Commit0) · rank 25 of 34
Unit
% resolved, higher is better
Configuration
Trinity Large Thinking (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
32.7%
Information gathering (GAIA) · rank 30 of 34
Unit
% resolved, higher is better
Configuration
Trinity Large Thinking (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
56.8%
Issue resolution (SWE-Bench) · rank 33 of 34
Unit
% resolved, higher is better
Configuration
Trinity Large Thinking (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.
Reported by OpenHands Index
33.3%
Testing (SWT-Bench) · rank 33 of 34
Unit
% resolved, higher is better
Configuration
Trinity Large Thinking (openhands)
Measured
30 Jun 2026
Not shown
Not the model alone: the score is for the model inside the OpenHands agent, at the SDK version shown.

Compare Trinity Large Thinking with