EASYHUB JOURNAL / 2025-01
AI news in January 2025
10 stories, ordered by their original dates. Open a title for its summary, image and source link.
- IT之家
Janus-Pro separates visual encoders while unifying understanding and generation
Janus-Pro separates visual encoding paths for understanding and generation within a unified system. The release offers research assets for multimodal design; benchmark rankings should not be generalized to every image or prompt.
- IT之家
Open-R1 starts a community effort to reproduce reasoning-model training
Open-R1 aims to reproduce missing pieces of reasoning-model training, including datasets and code. At this stage it is an open research initiative, not a finished system proven equivalent to the model that inspired it.
- IT之家
Qwen2.5-VL expands document and video understanding across three sizes
Qwen2.5-VL offers several sizes with improvements in spatial and temporal understanding, documents and visual agents. Different deployment budgets can use different variants, but their quality and resource requirements need separate evaluation.
- IT之家
Baichuan Omni-1.5 opens multimodal understanding and spoken output
Baichuan’s release combines multimodal inputs with interleaved text and audio output. Claimed performance in specialized domains is a model evaluation result, not evidence of clinical reliability or suitability for unsupervised decisions.
- IT之家
SmolVLM adds 256M and 500M variants for constrained devices
SmolVLM adds 256M and 500M variants for image captioning, visual questions and document or chart understanding on constrained devices. They extend the earlier 2B model. Parameter count alone does not fix memory usage, which also depends on image resolution, context and the inference backend.
- IT之家
Youdao opens a 14B model focused on step-by-step explanations
Youdao’s model focuses on stepwise explanations and a smaller deployment footprint. Exposing reasoning steps makes solutions easier to inspect, but visible steps alone do not establish mathematical correctness or replace independent checking.
- Tencent
Hunyuan3D 2.0 separates geometry generation from texture synthesis
Hunyuan3D 2.0 generates a mesh from an image and then textures it, with the texturing stage also usable on existing meshes. Tencent releases inference code, checkpoints and an online creation interface. Generated assets can be exported to standard 3D formats for further editing, rendering and content production.
- IT之家
DeepSeek-R1 launches alongside a family of distilled models
DeepSeek releases its reasoning model together with multiple distilled variants for different resource budgets. Vendor comparisons with o1 concern specific evaluations; the largest model’s results do not automatically describe every smaller derivative.
- IT之家
MiniCPM-o 2.6 explores on-device multimodal interaction at 8B
MiniCPM-o 2.6 supports text, image, audio and video inputs with text and speech output. The report highlights on-device interaction demonstrations, while actual device requirements depend on precision, runtime and workload.
- IT之家
Tiangong 4.0 adds reasoning and multimodal experiences with Skyo
Kunlun rolls reasoning and multimodal experiences into its web and mobile products alongside Skyo voice interaction. The report’s free-access description records the offering at that time, not a guarantee about today’s subscription terms.