Compare / Trinity Large Thinking vs Gemini 3.1 Pro Preview
Trinity Large Thinking vs Gemini 3.1 Pro Preview
13 results from 2 sources that measured both models. Each row is in the source's own units; there is no overall winner.
| Benchmark | Metric | Trinity Large Thinking | Gemini 3.1 Pro Preview |
|---|---|---|---|
| Arena (formerly LMArena)reported | OverallArena rating | 1,371 | 1,472 |
| Arena (formerly LMArena)reported | Business, management and financeArena rating | 1,367 | 1,477 |
| Arena (formerly LMArena)reported | Creative writingArena rating | 1,332 | 1,480 |
| Arena (formerly LMArena)reported | Expert promptsArena rating | 1,401 | 1,509 |
| Arena (formerly LMArena)reported | Instruction followingArena rating | 1,357 | 1,479 |
| Arena (formerly LMArena)reported | OverallArena rating | 1,368 | 1,487 |
| Arena (formerly LMArena)reported | Writing, literature and languageArena rating | 1,343 | 1,480 |
| OpenHands Indexreported | Average% resolved | 32.1% | 60.6% |
| OpenHands Indexreported | Front end (SWE-Bench Multimodal)% resolved | 25.0% | 44.1% |
| OpenHands Indexreported | Greenfield (Commit0)% resolved | 12.5% | 25.0% |
| OpenHands Indexreported | Information gathering (GAIA)% resolved | 32.7% | 81.8% |
| OpenHands Indexreported | Issue resolution (SWE-Bench)% resolved | 56.8% | 76.8% |
| OpenHands Indexreported | Testing (SWT-Bench)% resolved | 33.3% | 75.1% |
Bold green marks the better value on that metric. Where a source reports ranges that overlap, the difference may not be meaningful; see the benchmark page for ranges.