Apple’s Depth Pro generates 3D depth maps from images

Apple has developed a new AI model called Depth Pro that could revolutionize how machines perceive 3D vision. Michael Nuñez reports for VentureBeat that Depth Pro can generate detailed 3D depth maps from single 2D images in a fraction of a second, without relying on traditional camera data. The system outperforms previous models in speed …

Read more

Black Forest Labs releases Flux 1.1 Pro

Black Forest Labs has released a new, faster text-to-image model called Flux 1.1 Pro, reports Carl Franzen for VentureBeat. According to an independent benchmark, the model outperforms other AI image generators in terms of visual quality and speed. It generates images six times faster than its predecessor, with improved image quality and accuracy. Flux 1.1 …

Read more

DeepMind’s SCoRE makes AI models more reliable

DeepMind has developed a new technique called SCoRe that significantly improves the self-correction abilities of large language models (LLMs). Ben Dickson reports this in an article for VentureBeat. SCoRe uses self-generated data and enables LLMs to use their internal knowledge to identify and correct errors. In tests, SCoRe significantly outperformed other self-correction methods. The technique …

Read more

Nvidia surprises with powerful, open AI models

Nvidia has released a powerful open-source AI model that rivals proprietary systems from industry leaders like OpenAI and Google. The model, called NVLM 1.0, demonstrates exceptional performance in vision and language tasks while also enhancing text-only capabilities. Michael Nuñez reports on this development for VentureBeat. The main model, NVLM-D-72B, with 72 billion parameters, can process …

Read more

Mostly AI helps companies to train AI without privacy concerns

Mostly AI has introduced a new feature for generating synthetic texts. The tool allows companies to use confidential data such as emails and conversations for AI training without privacy concerns. As reported by Shubham Sharma, the platform generates a version of proprietary information free from personally identifiable data. This enables businesses to train and optimize …

Read more

OpenAI’s DevDay focussed on optimizations

OpenAI unveiled several new features at its DevDay 2024 developer conference, aimed at making AI applications more accessible and affordable. The focus is on the Realtime API for real-time speech applications, Vision Fine-Tuning to enhance visual AI capabilities, Model Distillation for optimizing smaller models, and Prompt Caching for cost savings. These innovations are designed to …

Read more

Together AI introduces new platform for private AI

Together AI has unveiled a new platform for AI applications in private cloud environments. The system promises faster inference and lower costs for enterprises. CEO Vipul Prakash told VentureBeat that the platform can increase AI inference performance by two to three times and cut hardware requirements in half. A key feature is the flexible orchestration …

Read more

Molmo to improve AI agents

A new open-source AI model called Molmo could help advance the development of AI agents. Developed by the Allen Institute for AI (Ai2), the model can interpret images and communicate via a chat interface. According to Wired’s Will Knight, this enables AI agents to perform tasks such as web browsing or document creation. In some …

Read more

Scramble aims to become a Grammarly alternative

The AI tool Scramble is an extension for the Chrome browser. Once installed, you highlight the text in question, select “Scramble” from the context menu, and get suggestions for improvement. According to the project’s official GitHub page, it aims to be a more flexible and privacy-friendly alternative to Grammarly. However, the privacy argument in particular …

Read more

EzAudio creates high quality sound effects

Researchers at Johns Hopkins University and Tencent AI Lab have developed a new text-to-audio model called EzAudio. As Michael Nuñez reports for VentureBeat, EzAudio can generate high-quality sound effects from text descriptions. The model uses an innovative method for processing audio data and a new architecture called EzAudio-DiT. In tests, EzAudio outperformed existing open-source models …

Read more

×