EASYHUB JOURNAL / 2024-11
AI news in November 2024
11 stories, ordered by their original dates. Open a title for its summary, image and source link.
- IT之家
Mooncake opens components for KV-cache-centered disaggregated inference
Mooncake advances a cache-centered inference architecture with separated prefill and decode resources. The release is staged, starting with components such as its transfer engine, rather than completing every planned storage and serving integration at once.
- IT之家
SmolVLM targets low-memory visual-language inference at 2B parameters
SmolVLM reduces visual-language inference costs with a compact backbone and compressed image representations. Base and tuned variants accompany open training assets, offering a smaller starting point for device-specific experimentation and fine-tuning.
- 量子位
aisuite offers one calling interface across model providers
aisuite wraps multiple model providers behind a consistent interface, reducing integration boilerplate when switching services. Developers still need the relevant credentials, dependencies and service access; the wrapper does not make paid inference free.
- 量子位
OpenScholar combines retrieval and an 8B model for literature synthesis
OpenScholar combines a scholarly retrieval collection, reranking and a specialized 8B model to synthesize literature. Its open research assets make the pipeline inspectable, while generated citations and conclusions still require checking against the papers.
- 新智元
Spirit LM studies interleaved speech and text with expressive style
This research report explains Spirit LM’s interleaved text-and-speech training and its base and expressive variants. The work targets cross-modal language generation; expressive demonstrations do not establish production reliability across all voices and languages.
- IT之家
Tencent ima connects public-account search, local material and notes on desktop
Tencent presents ima’s Windows and Mac workspace, combining web search including WeChat public-account articles with local document reading and notes. Users can ask questions, extract summaries, create mind maps, and continue writing or translating in the same workspace, reducing switching between research, reading and recording information.
- IT之家
Qwen2.5-Coder expands its size range for different coding workloads
Qwen expands its coding family across six sizes, covering lightweight deployments and more demanding generation tasks. Available checkpoints and demos support evaluation, but results for the largest model should not be attributed to every smaller variant.
- 太平洋科技
Baidu previews Miaoda’s multi-agent approach to no-code applications
Baidu presents Miaoda at its 2024 conference, combining natural-language requirements, multiple agents and tool calls for building applications. The story records an initial demonstration and product direction. The cited report does not establish general public availability, so the announcement is not described as an unrestricted public beta.
- IT之家
Meta MobileLLM explores compact language models for phones
The report covers MobileLLM’s release and expanded model sizes for resource-constrained devices. Its significance is compact architecture and deployment potential, not identical performance across every phone or checkpoint.
- IT之家
Hunyuan3D-1.0 opens text- and image-conditioned 3D generation
Hunyuan3D-1.0 generates 3D assets from text or a single image through two stages: multi-view image generation followed by geometry reconstruction. Tencent opens code, weights and research material for further modeling workflows. Reported stage timings describe the demonstrated setup, not identical performance on every device.
- IT之家
HybridFlow and veRL open a flexible reinforcement-learning training stack
ByteDance and the University of Hong Kong introduce HybridFlow, released as veRL. The framework combines control approaches for reinforcement-learning training and inference; reported throughput improvements are experimental results rather than universal speed guarantees.