What is HappyHorse?
On April 8th, a model named HappyHorse received a high score on the well-known AI video model leaderboard. This ranking is called Artificial Analysis. In the ranking, the term "anonymous submission" was used to label the HappyHorse model, indicating that no one knows who submitted this model for now. However, its score is a full 60 points higher than ByteDance's Seesaw 2.0 video model.
But in the AI era, the scores on the rankings have actually become numbingly routine. Many new models, once released, are sure to claim that they score much higher than other similar models on certain benchmarks. But this time it's different because the Artificial Analysis ranking is quite different from many other AI rankings. Many other AI rankings might allow you to boost your scores by practicing corresponding test questions in a lab setting. But this ranking uses real user blind testing to score you. There's another very interesting topic here, which is that this type of blind test ranking actually has high commercial value because it can provide model manufacturers with very real feedback, which they can use to improve their own models. This is a side note, meaning that such rankings are still valuable. But if we want to launch a project like this, there will be a challenge: how to get users to spend time testing the model and scoring it for you. This is what needs to be designed in the business model.
In short, you cannot achieve a good ranking on the Artificial Analysis chart thru manipulation or cheating. Users just watch two videos, and whichever one they prefer, they vote for it. Ah, so on this list, HappyHorse's high score is quite valuable.
First, let's talk about the specific data. HappyHorse scored 1330 points on the tattoo video track, which has no audio, and 1392 points on the image generation video track. So, in these two categories, it is currently ranked number one in the world. Ah, of course, I just emphasized that it is audio-free. Text-to-video generation track. For video generation with audio, C Station 2.0 still leads Happy House, with HappyHorse in second place. C Station 2.0 is currently 14 points ahead, which is a very small margin. Ah, this somewhat indicates that HappyHorse is extremely strong in image generation. If we ask users to judge images without audio, HappyHorse is leading by a wide margin. As for the audio and video part, it doesn't perform as well, which indicates that there is still significant room for improvement in having the HappyHorse model generate both video and audio simultaneously.
As soon as this model was released, API pass tried to integrate with it. Ah, but we haven't found any public APIs to integrate with, so we might need to wait a bit longer. But the API pass team is closely monitoring all information about this model, and as soon as it goes live, we will integrate it immediately.
What are HappyHorse's strengths?
The next point we might consider is that since Seedance 2.0 has been out for a while, ByteDance has a strong presence in the video field due to TikTok, which has a significant data accumulation. Therefore, people believe that their video model should be the best. Another regretful example is China's Kuaishou, specifically their Keling model. Actually, Kuaishou also has a lot of video data, but the unfortunate thing is that when we look back, we find that Kuaishou was initially leading in many areas, but it was quickly surpassed. Now, the discussions about Kuaishou have decreased. A while ago, everyone was discussing seedance 2.0 more, and then seedance 2.0 opened up its API, but you need a very high amount of money, for example, you need to spend 10 million to get access to this API. The awkward part is that we just got this permission not long ago, and then HappyHorse came out. Ah, so who gets hurt in this situation? The ones who are hurt are us, the ones who spent 10 million to get the seedance 2.0 interface, and now its performance might not even be the best one.
Let's go back and discuss HappyHorse. From the blind test cases published on the leaderboard, we can see that HappyHorse has several outstanding aspects. First, its cinematography is more natural. Here, "natural" means that with the same prompt, the images generated by C Station might have the characters standing in the middle, with little camera movement. And HappyHorse feels more like having a real person carrying the camera, using slow zoom-ins, panoramic transitions, and changes in light and shadow to create a stronger sense of cinematic language.
Another point to note is that the continuity of character movements will be better. We all know that when generating videos, we humans are very sensitive to judging our own kind. For example, if we generate a video with a character in motion, if the character's movement is not smooth, we can quickly tell and feel that the video looks a bit fake. Usually, when we encounter videos like this, we might choose a slightly better one by generating it multiple times. However, in this situation, HappyHorse generates better quality. In other words, you might be able to reduce the number of generations and still obtain a higher quality video. In many models, such as the Keling 3.0 model, during the movement of characters, in the latter half of the video, as I mentioned earlier, they tend to deform easily. And HappyHorse is relatively stable in these metrics. A 5-second video, you can hardly tell. Ah, it's AI-generated.
Additionally, some information shows that HappyHorse generates faster. Ah, many people might deeply feel the fast generation speed because when we use Sendance 2.0 or other video models, the waiting time is really too long. Waiting for a long time will affect our experience with this tool because if it's too slow, it's inefficient. And HappyHorse can generate a 1080P, 5-second video on a single H100 GPU in just 38 seconds. Under the same conditions, sendance 2.0 might take 60 seconds, making it almost twice as fast.
Now, two or three days after the model's release, the team's identity has gradually been revealed. They come from Alibaba's Taotian Group's Future Life Lab, and the head of this lab is named Zhang Di. Zhang Di is quite a well-known figure in the field of video generation. He himself works in big data and machine learning, and his professional experience is mainly related to video. He previously worked on Kuaishou's Keling. The underlying architecture of the large models, Kuaishou's 1.0 and 2.0, were both developed by him. At one point, it was the benchmark product for AI generation in China. Some in the industry also call him the father of Keling. Then, he briefly joined Bilibili as the technical lead, but left after a month. He returned to Alibaba and then went to the Future Life Lab to work on the HappyHorse product.
Is it open source?
Whether HappyHorse will be open-sourced is actually a very important matter for the entire AI video generation field. On April 8th, we observed that the corresponding links on GitHub and Hugging Face both displayed "coming soon," and the official website also claimed that the model weights, distilled versions, super-resolution modules, and inference code would all be open-sourced.
Finally, when you use Google to search for HappyHorse, you will find many websites or open-source repositories related to HappyHorse. Most of these websites and open-source repositories have nothing to do with the official HappyHorse; they are all phishing sites. Before the official code is released and open-sourced, any website claiming to offer an online experience of HappyHorse should be approached with caution.
