EASYHUB JOURNAL / 2025-08
AI news in August 2025
12 stories, ordered by their original dates. Open a title for its summary, image and source link.
- StepFun / 阶跃星辰
Step-Audio 2 mini opens weights for audio understanding and conversation
StepFun opens Step-Audio 2 mini and Base checkpoints with inference examples for audio understanding and spoken interaction. Tool-assisted retrieval can complement responses. Developers can run the models locally, while the official online demonstration requires a platform API key; the hosted service and downloadable weights are separate offerings.
- IT之家
IBM and NASA open Surya for solar-observation modeling
Surya models solar observations and opens research assets for space-weather studies. It can support development of forecasting methods without replacing operational monitoring or warning procedures.
- IT之家
VibeVoice-1.5B explores long-form multi-speaker speech generation
VibeVoice targets long-form spoken content with multiple speakers, while the report notes language and overlapping-speech limits. Reference voices require permission and synthetic output should be disclosed rather than used for impersonation.
- Google
NotebookLM expands Video Overviews beyond 80 languages
NotebookLM expands Video Overviews beyond English to more than 80 languages and makes non-English Audio Overviews more detailed. Shorter audio summaries remain available. The update broadens language coverage and explanation depth; video overviews still present the notebook's supplied materials rather than inventing unrelated footage.
- IT之家
CodeBuddy IDE enters public testing with design-to-code and deployment tools
CodeBuddy IDE’s Chinese edition enters public testing without invitations, using DeepSeek-V3.1 at announcement. Beyond conversational coding, it supports design-to-code workflows and connects CloudBase, EdgeOne Pages and Supabase for backend and deployment tasks such as databases, authentication and site previews.
- IT之家
Fun-ASR adds industry vocabulary to DingTalk meeting transcription
DingTalk and Tongyi integrate Fun-ASR into subtitles, interpretation, meeting notes and voice assistance, with industry vocabulary and enterprise customization. Using company contacts, calendars or knowledge requires authorization. This product integration is not proof that the entire service is freely deployable from a similarly named open-source project.
- 量子位
Nemotron Nano v2 pairs a compact model with broader training assets
The report covers NVIDIA’s compact reasoning model and expanded access to training assets. Throughput comparisons depend on hardware and inference settings rather than applying uniformly to every workload.
- IT之家
VeOmni separates model computation from distributed training orchestration
VeOmni offers a PyTorch-native framework for composing parallelism around multimodal models. Reported engineering and throughput gains describe particular experiments, while adoption requires validation on the target cluster and architecture.
- Vercel
v0.app expands from UI generation to planning and debugging applications
Vercel moves v0.dev to v0.app and expands the builder with planning, web research, file reading and debugging. The workflow combines interfaces, content and backend logic, allowing users to refine an existing project through conversation instead of repeatedly starting over with isolated code-generation prompts.
- IT之家
MiniCPM-V 4.0 pairs compact vision modeling with an open iOS app
MiniCPM-V 4.0 combines a compact visual-language model with iPhone and iPad application examples. Device-specific latency demonstrations are configuration-dependent rather than guarantees for every handset.
- IT之家
Huawei announces a broader opening of the CANN development ecosystem
Huawei outlines broader access to CANN and related tooling for accelerator and application development. The announcement concerns ecosystem strategy; availability and licenses still need checking at the component and version level.
- IT之家
Qwen-Image opens image generation with stronger bilingual text rendering
Qwen-Image emphasizes complex text rendering and offers model assets, code and demonstrations. Poster lettering and numerical details still need inspection; benchmark performance cannot replace checking the generated image.