Yes, I did make a Mistral account. Forgot to mention that I tested all of the mentioned chatbots on free plans, except Copilot, which I had the highest tier plan for, as part of my job. Needless to say, even on the highest tier Copilot performed the worst. Congratulations, Microsoft.
I do appreciate Mistral being European and maintaining better values than American competition. That’s why I feel bad rating it so low. Although as you said, many tasks that LLMs are useful for don’t require the cutting edge at all.
In fact, Mistral Vibe (I preferred the name Le Chat, but oh well) is probably more than enough for most people. It’s just that availability of much better models makes it look worse than it is.
There’s also the question of system prompts. ChatGPT has been known to work considerably better with a good prompt setup, to the point where it’s not one of the worst chatbots anymore. I was only testing the defaults. It’s entirely possible that setting the right system prompt for Mistral improves it noticeably.
At the end of the day, however, if I can run a small, open-weight model locally with comparable results - using a huge, cloud-based and closed model makes no sense to me. I’m not connecting to any cloud unless I really need to.
Alibaba is releasing Qwen 3.8 27B in a little over 3 hours from now, so I’ll have to try that too. Though I’ll probably wait for fine-tuned community releases to make a more detailed comparison.