Researchers developed ML3DHS, an AI framework that assigns multiple related labels to 3D points, enabling machines to understand objects at different levels of detail. This improves navigation, interaction, and perception for robots, autonomous systems, and interactive 3D technologies.
Researchers developed the Motion Style Slider framework for continuous control of motion style intensity in generated human animation. The framework achieves consistent style control and smooth transitions using only two motion samples, allowing animators to refine character performances according to their creative vision.
The robotic lab can autonomously assemble and fine-tune optics experiments, reducing manual setup time from days or months to minutes. This could enable scientists to focus on theoretical work and accelerate innovation in fields like quantum technologies and renewable energy.
Researchers developed xvr, a patient-specific AI technique that accurately matches X-rays with 3D medical scans, improving surgical navigation and safety. This innovation enables faster and more precise minimally invasive surgeries, particularly in fields like orthopedics and neurosurgery.
A two-level AI framework is developed to identify the presence and extent of visible corrosion and classify corrosion pixels into four visual categories. This framework provides complementary information about where corrosion occurs and how accurately its boundaries are represented.
Researchers from SUTD develop HieraScaffold, an AI framework that generates large-scale 4D LiDAR scenes more efficiently and coherently. The framework captures both static structures and moving objects, improving the realism and accuracy of autonomous systems.
Researchers develop a novel framework, LL-Refiner, to enhance high-resolution images in poor lighting conditions, outperforming state-of-the-art techniques. The framework uses a coarse enhancement stage to guide the recovery of fine details, resulting in improved visual quality and performance in downstream computer-vision tasks.
Researchers developed a system using a 3D time-of-flight camera to measure frozen skipjack tuna, capturing dense 3D point-cloud data for accurate body size and weight estimation. The combination of 3D imaging and machine learning analysis yielded accurate noncontact estimates of fish body weight.
A novel AI model has been developed that can recognize yoga poses with high accuracy, paving the way for more effective digital coaching tools and movement-monitoring applications. The model achieved accuracy levels of over 93% during testing, significantly outperforming previous models.
Researchers from NUS CDE have developed a computational framework that reconstructs the genealogy of vernacular architecture, including Singapore's historic shophouses. The framework generates a detailed architectural 'family network' that reveals how different styles evolved and interacted over time.
Hyunsoo Lee, an SNU undergraduate, presents research in generative visual computing at leading conferences NeurIPS, CVPR, and ECCV. His work spans image editing, human motion, and 3D content generation, leveraging pretrained generative models to produce consistent outputs.
The new photonic architecture harnesses three fundamental degrees of freedom: wavelength, mode, and polarization, achieving 192 parallel computing channels. The chip supports large, reconfigurable convolution kernels up to 13x13, capturing global structural contours while preserving fine details.
University of Jaume I students Pau Montagut and Mario García won first place at the ICRA 2026 robotics conference with an AI model that taught a Toyota HSR robot to perform household tasks. The team's achievement is notable given their undergraduate status, competing against teams of research personnel.
A new underwater mapping technique, Sonar-MASt3R, combines sonar and visual data to generate detailed 3D maps of environments in real-time. The system enables vehicles to navigate through cloudy water by quickly mapping the general shape of their surroundings using sonar.
Researchers at Penn State developed photomemristors that adjust sensitivity based on light levels, like the human eye. These devices can process light data faster and more accurately than traditional systems in mixed lighting environments.
A new ultra-lightweight AI model, Multinex, advances low-light image enhancement by leveraging classical colour vision theory and Retinex principles. The model outperforms comparable compact systems, recovering detail and clarity from previously unusable images.
Researchers developed SpiderCam, a highly energy-efficient 3D camera inspired by jumping spiders. It captures two images with different focus settings and analyzes the differences in sharpness to produce real-time 3D maps while consuming less than a watt of power.
Researchers from MIT and IBM create a state-of-the-art dataset called ChartNet, which includes over a million varied charts. The dataset is designed to teach vision-language models how to effectively interpret charts, enabling them to outperform commercial models on tasks like data extraction and chart summarization.
Researchers from Drexel University developed BioCoach, a program using AI and computer vision to analyze video and provide form coaching in real time. The system analyzes visual appearance and motion patterns, as well as 3D skeletal movements and body shape, to deliver detailed biomechanics-based feedback.
Researchers developed a disco laser system to enhance data visualization for snow groomers, improving operator comfort and reducing nausea caused by VR headsets. The system also enables better tracking and orientation aids, leading to more efficient and safe operation in challenging conditions.
Brown University computer scientists introduce PackUV, a compression method that enables everyday video formats to stream volumetric video. The technique improves capture of 3D action and makes final products compatible with existing video codecs.
A new approach enables computers and machines to capture images at higher resolution and faster speed, making it impervious to reflective surfaces. The technology uses a virtual screen created by repurposing the surroundings of specular objects.
A systematic review of AI models for meningioma segmentation reveals that better model architecture is the key driver of improved performance. The top models achieved high accuracy and efficiency, while future research focuses on making them more generalizable and efficient for real-world clinical settings.
Machines with advanced AI capabilities are being developed to restore lost senses such as sight and sound, and even simulate touch and taste. This technology has the potential to revolutionize various fields like healthcare and education, but also raises serious concerns about privacy and manipulation.
A new technique uses a single image to forecast solar panel energy production and maximize output. The method estimates the amount of energy that will be produced based on the angle of the sun, shadows, reflections, and weather patterns, allowing for more accurate placement and optimization of solar panels in urban areas.
A team of scientists from NTU Singapore has developed a new biochip that, when paired with Artificial Intelligence (AI), can detect quickly and accurately extremely small amounts of microRNAs. The device can cut detection time from hours to 20 minutes.
HeapGrasp uses RGB images to analyze object silhouettes and estimate its 3D shape, reducing the need for depth information. The approach achieves high accuracy while minimizing camera movement and execution time.
A new model combines multiple ways of analysing 3D data, integrating local and global perspectives to interpret complex environments more reliably. The system improves detection of small or partially visible objects in real-world situations, enhancing safety in autonomous systems.
Researchers at MIT developed a new method that coaxes AI models to achieve better accuracy and clearer explanations in safety-critical applications. The approach extracts concepts the model has learned while training for a specific task and forces it to use those, producing better explanations than standard concept bottleneck models.
Researchers used machine learning techniques to compress a large model of the visual cortex, creating smaller versions that predict neural responses with high accuracy. The compact models revealed specific computational patterns in how neurons detect important features, offering insights into how visual information is processed.
Drexel researchers create machine learning program that integrates qualitative and quantitative data to identify gentrification in Philadelphia neighborhoods. The program, trained with data from thousands of images and focus groups, accurately identifies new-build gentrification with 84% accuracy.
InstaDrive generates precise editing of vehicles and map elements, enabling efficient labeled data generation. It outperforms baselines in FID and mAP, preserving accurate map structures and maintaining multi-view consistency.
A team of MIT engineers developed a deep-learning model that predicts how individual cells will fold, divide, and rearrange during a fruit fly's earliest stage of growth. The model achieved 90% accuracy in predicting the movement of 5,000 cells over the first hour of development.
A breakthrough AI system called OmniPredict can predict human pedestrian behaviors with unprecedented accuracy, revolutionizing self-driving cars and urban mobility. The model combines visual cues with contextual information to anticipate pedestrians' next moves, reducing the risk of accidents and improving traffic safety.
A UBC Okanagan team harnesses computer modeling to study wildfire movement, finding that fires often behave randomly due to factors like fuel type, wind, and terrain. This randomness can lead to significant variations in fire spread, highlighting the need for more probabilistic models.
Researchers at Purdue University are testing a computer-vision method to analyze smartphone photos of pregnant women's eyes to predict preeclampsia risk. The two-year study aims to reduce maternal mortality in Africa and could potentially save thousands of lives.
Researchers developed AI-powered BlinkWise glasses that track blinking patterns to assess fatigue, mental workload, and eye-related health issues. The device uses radio signals to detect minute eyelid movements with unprecedented detail, preserving privacy and using minimal power.
Researchers at Purdue University have developed an algorithm that recovers detailed spectral information from photographs taken by conventional cameras. The method uses computer vision, color science, and optical spectroscopy to achieve high spectral resolution comparable to scientific spectrometers.
Researchers at UMC Utrecht developed a new AI-powered printer called GRACE that can print implantable tissues with improved cell survival and functionality. The printer uses computer vision and laser-based imaging to design and print complex structures, including blood vessels and cartilage layers.
Researchers at Brown University developed an image processing technique that harnesses camera motion to increase resolution, producing super-resolution images with details sharper than the original pixel array allows. The technique has potential applications in archival photography and photography from moving aircraft.
A research team developed an innovative unsupervised model for industrial anomaly detection using paired well-lit and low-light images. The model leverages feature maps, Low-pass Feature Enhancement, and Illumination-aware Feature Enhancement to detect anomalies while remaining lightweight and memory-efficient.
A new study reveals that pedestrians are now walking faster and spending less time in public spaces. Researchers analyzed 40 years of video footage to find a 14% decline in people lingering in these areas.
Researchers developed CoSyn, a new approach to train open-source models using AI-generated scientific figures and charts. The resulting dataset, CoSyn-400K, includes over 400,000 synthetic images and 2.7 million sets of corresponding instructions. CoSyn-trained models match or outperform proprietary peers in various benchmark tests.
MIT engineers developed a versatile demonstration interface that allows users to teach robots new skills in three intuitive ways: remote control, physical manipulation, or demonstration. This innovation expands the type of users and 'teachers' who interact with robots, enabling robots to learn a wider set of skills.
Researchers have demonstrated a new technique, RisingAttacK, to manipulate all widely used AI computer vision systems, allowing them to control what the AI 'sees'. The attack is effective at influencing the AI's ability to detect top targets, such as cars, pedestrians, or stop signs.
A new study reveals a five-fold increase in computer vision papers linked to surveillance patents, highlighting the rise of obfuscating language that normalises surveillance. The top institutions producing surveillance are Microsoft, Carnegie Mellon University, and MIT.
Researchers at UMass Amherst created integrated arrays of gate-tunable silicon photodetectors that can capture dynamic visual information and classify static images with high accuracy. The technology has the potential to reduce latency in computer vision tasks, enabling applications like self-driving vehicles and bioimaging.
Researchers at MIT developed a simulation method that allows for accurate and stable simulations of elastic materials, enabling the creation of realistic bouncy characters in movies and video games. The approach preserves physical properties and avoids instability, making it a promising tool for engineers to design flexible products.
Researchers have developed an image-analysis tool called SeaSplat that cuts through the ocean's optical effects and generates images of underwater environments with accurate colors. The team paired SeaSplat with a computational model to convert images into three-dimensional underwater worlds, allowing for virtual exploration.
Researchers developed an innovative deep-learning-based framework that uses common surveillance cameras to estimate rainfall in real time. The approach achieved high predictive accuracy across various environmental conditions and lighting scenarios, outperforming traditional methods while maintaining low computational costs.
A new deep learning model, ENDNet, significantly enhances subgraph matching accuracy by identifying and neutralizing extra nodes that interfere with the matching process. This improves performance in pattern recognition tasks across various fields, including drug discovery and natural language processing.
Researchers at MIT developed a technique to improve the reliability of conformal classification, which can produce impractably large prediction sets. By combining test-time augmentation with conformal prediction, they reduced prediction set sizes by up to 30 percent while maintaining probability guarantees.
Schmid's contributions have helped computers recognize complex objects, understand video analysis, and process realistic settings. Her leadership has built active research communities, mentoring and supervising peers across the field of computer vision.
A University of Florida researcher has developed a groundbreaking AI tool called VisionMD that analyzes videos of patients with Parkinson's disease and other movement disorders. The tool provides valuable information about how the disease is progressing and responding to medications, improving patient care and advancing clinical research.
A collaborative research team has developed a novel mixed reality (MR) technology that uses real-world doors as natural transition points. The system allows users to select a door within their MR interface and seamlessly transition into a virtual space, creating an unprecedented sense of immersion.
Researchers at the University of Arizona have developed a new 3D imaging technique, deflectometry, paired with advanced computation to improve eye-tracking accuracy. The method can capture gaze direction information from more than 40,000 surface points, theoretically millions, increasing accuracy by a factor of over 3,000 compared to c...
Researchers develop a new approach combining Phase Measuring Deflectometry and Shape from Polarization to accurately image specular surfaces without prior knowledge or assumptions. The single-shot method enables motion-robust measurements, pushing the limits for next-generation 3D sensors.
Researchers have developed a hybrid image-generation tool called HART that combines the strengths of autoregressive and diffusion models. It achieves high reconstruction quality with significantly reduced computational resources, enabling local execution on laptops or smartphones.
The conference aims to bridge theoretical advancements with practical applications in AI and visual computing. Researchers can submit original research papers and attend keynote sessions, offering opportunities to network with pioneers in intelligent technologies.
Scientists developed a method that harnesses chromatic aberration to produce high-quality images using a single exposure. The AI approach uses generative models to retrieve phase information from limited data input.