EASYHUB JOURNAL / 2025-03
AI news in March 2025
10 stories, ordered by their original dates. Open a title for its summary, image and source link.
- IT之家
AutoGLM Rumination previews research combined with action
Zhipu previews an agent combining research, planning and tool execution. The announcement distinguishes its current research-oriented preview from later action capabilities and planned open releases.
- IT之家
Wenxiaoyan combines model routing, voice and image questions
Baidu updates its assistant with automatic model routing, voice interaction and image queries. The product combines multiple systems, illustrating application orchestration rather than every feature belonging to one underlying model.
- IT之家
Qwen2.5-Omni combines multimodal input with streaming speech
The Thinker-Talker design links multimodal understanding with text and speech responses. Real-time interaction depends on the full capture, transport and inference pipeline, not only the released model.
- IT之家
Hunyuan T1 launches with stronger reasoning over code and long documents
Hunyuan T1 uses reinforcement learning for mathematics, logic, science and coding, with the Mamba-Transformer architecture inherited from Turbo S to reduce long-context reasoning overhead. A web interface and Tencent Cloud API are available at launch for document analysis and questions requiring multi-step reasoning.
- IT之家
Llama Nemotron separates edge, single-GPU and server deployment tiers
NVIDIA describes reasoning-model tiers and visual-language work for agent deployments. Family announcements cover different availability stages, making checkpoint-level and service-level verification important.
- IT之家
Mistral Small 3.1 targets local multimodal and low-latency use
The report covers local deployment, longer contexts and tool use in Mistral Small 3.1. Hardware examples depend on precision and context settings, so a device model name alone cannot guarantee the reported experience.
- The Verge
Gemini Canvas adds real-time document and code editing
Gemini adds a workspace for editing documents and previewing code, alongside audio summaries. At launch, the audio feature was English-only, with additional languages planned rather than already available.
- 量子位
OpenManus combines browser, code and file tools in a modular agent
OpenManus combines planning, browser actions and file tools using an existing open-source ecosystem. The report describes a modifiable prototype approach, not proof of full parity with a commercial agent’s reliability.
- IT之家
Tencent opens a 13B image-to-video model with LoRA training code
Hunyuan expands into image-to-video generation using a reference image and motion or camera instructions. Tencent releases the 13B model weights, inference code and LoRA training code. The online service also showcases lip-sync, motion driving and audio features, which should not all be attributed to the single image-to-video checkpoint.
- IT之家
QwQ-32B opens a reinforcement-learning-based reasoning model
QwQ-32B adds an openly available reasoning model with online access. Its reported mathematics, coding and tool-use results are benchmark-specific; parameter count alone does not determine end-to-end task cost.