EASYHUB JOURNAL / 2025-02
AI news in February 2025
12 stories, ordered by their original dates. Open a title for its summary, image and source link.
- IT之家
Hunyuan Turbo S introduces a hybrid architecture for faster responses
Tencent introduces Turbo S, a fast-response model combining Mamba and Transformer components to reduce long-context compute and cache costs. Cloud API access is available at announcement, with Yuanbao integration to follow. The model also serves as a foundation for Hunyuan reasoning, long-document and coding variants.
- IT之家
Phi-4 adds mini and multimodal variants
Microsoft expands Phi-4 with a smaller language variant and a separate multimodal model. Supported inputs, outputs and serving requirements differ, so family-level claims should not be assigned to every checkpoint.
- IT之家
Magma explores agents spanning digital and physical tasks
Microsoft’s research model explores shared representations for screen-based and physical actions. Demonstrations do not guarantee safe shopping or robotics behavior; authorization and execution checks remain necessary.
- IT之家
Wan 2.1 opens video generation with smaller and larger model options
Wan 2.1 releases video weights and inference code across different model sizes. Consumer-GPU demonstrations apply to particular settings and resolutions, not to the memory requirements of every variant.
- IT之家
Moonlight and distributed Muon open a route to training-efficiency research
Moonshot shares MoE checkpoints, optimizer code and intermediate training assets around Muon. The contribution concerns inspectable training methods; total model size and parameters activated per step are distinct measurements.
- IT之家
SkyReels releases tools for character video and performance control
SkyReels combines character-oriented video generation with methods for controlling expressions and movement. Text, images and driving footage support different workflows, while final clips still require consistency checks and editing.
- IT之家
StepFun and Geely open video and voice-interaction model assets
The joint release covers separate systems for video generation and controllable voice interaction. These are different models and serving pipelines, so their capabilities and deployment requirements should be assessed independently.
- Dify
Dify 1.0 expands application workflows through plugins and agent nodes
Dify 1.0 introduces an extensible plugin architecture and agent nodes for combining models, tools and decision logic in application workflows. Marketplace and custom integrations expand its reach. Plugins can access real services and data, so their provenance and granted permissions still require review.
- IT之家
OmniParser V2 structures screen elements for computer-use agents
OmniParser V2 structures buttons and other screenshot elements for language-model action planning, improving small-target detection. OmniTool supplies a Dockerized Windows experimentation environment linking perception and execution. The parser alone is not a complete agent, and the research guidance recommends permission boundaries and human oversight.
- 新智元
OpenThinker-32B shares reasoning weights and training data
The OpenThoughts effort shares a reasoning model, training data and generation pipeline. Reported benchmark results concern a specific setup; the open assets make data choices and subsequent experiments more inspectable.
- IT之家
VideoWorld explores learning knowledge and actions from video
VideoWorld studies learning from visual sequences without centering the pipeline on a language model. Its game and robotics experiments illustrate task-specific potential rather than a general solution to understanding the physical world.
- 机器之心
FireRedASR opens alternative architectures for speech recognition
Xiaohongshu’s FireRed team releases models and inference code using encoder-decoder and LLM-integrated approaches to Mandarin transcription. They target different efficiency and accuracy needs. FireRedASR recognizes speech rather than generating voices, and reported benchmark rankings remain specific to their evaluation datasets.