← All news Models

SenseTime released the multimodal model SenseNova-U1.5-8B-MoT

September 15, 2026
SenseTime released the multimodal model SenseNova-U1.5-8B-MoT

Company SenseNova-U1.5-8B-MoT-SFT introduced a new architecture designed for working with visual content. The model is based on the Mixture-of-Tokens (MoT) approach and optimized for performing image generation and editing tasks.

Developers implemented support for output in 4K resolution, which sets the system apart from analogs with a comparable number of parameters. The model has 8 billion parameters and underwent the SFT (Supervised Fine-Tuning) stage to improve the accuracy of following user instructions when creating graphics.

The technical implementation of the model allows it to be used as a native multimodal system capable of processing complex visual queries. The model's weights and deployment documentation are already available for download in open access on the Hugging Face platform.

The integration of the MoT architecture allows for more efficient distribution of computational resources when processing visual tokens. This solution is aimed at reducing delays when generating high-resolution images while maintaining a compact model size for local launch on specialized equipment.

The new release expands the ecosystem of tools for working with generative content, offering developers a flexible tool for embedding image editing functions into third-party applications. The source code and model weights are distributed for research and applied purposes.

mozgi.io — AI tools and prompts
Copied