Outline
New customer case study in the works with LMArena !
LMArena evaluates LLMs through side-by-side, blind comparisons. Users see two anonymous responses to the same prompt and vote on which one is stronger.
Concept
Important: LMArena is doing a rebrand at the end of January and will send us a new logo. I'm not sure what else will change about their branding but will keep you posted as I learn more)
For the image, something with an AI battle / single combat / gladiator vibe would be fun, but I'm cool with wherever you want to take this.
Here's a few key components of their product if it's helpful:
- There's a leaderboard for each modality (text, video, code...)
- The comparison battles are always 1 on 1
- The AI outputs are always anonymous until you vote (to prevent biased voting)
- Human preference data is actually really useful to frontier labs (OpenAI, Meta, Anthropic...) to benchmark models and test unreleased variants
Deadline & Priority
We're planning to publish the first week of February, so end of January would be much appreciated!
Technical requirements
Standard case study image size
Outline
New customer case study in the works with LMArena !
LMArena evaluates LLMs through side-by-side, blind comparisons. Users see two anonymous responses to the same prompt and vote on which one is stronger.
Concept
Important: LMArena is doing a rebrand at the end of January and will send us a new logo. I'm not sure what else will change about their branding but will keep you posted as I learn more)
For the image, something with an AI battle / single combat / gladiator vibe would be fun, but I'm cool with wherever you want to take this.
Here's a few key components of their product if it's helpful:
Deadline & Priority
We're planning to publish the first week of February, so end of January would be much appreciated!
Technical requirements
Standard case study image size