60 Percent of Enterprise Apps Now Use Multimodal AI. LG's Washing Machine Still Just Says AI Mode.

60 Percent of Enterprise Apps Now Use Multimodal AI. LG's Washing Machine Still Just Says AI Mode.

Independent Analysis · Physical AI
Nearly 60 percent of enterprise applications now run on models that combine two or more data modalities. A two-minute photo diagnosis on a broken garage door sensor shows what that actually looks like in practice, and why LG's "AI Mode" and appliance makers like it are not shipping it yet.
By Shashi Bellamkonda
Illustration of a house outline with a home assistant interface and thermostat
2 min
from photo to fixed garage door
0
appliances in my house that predicted a failure before it happened

The garage door stopped closing halfway last night. It kept reversing itself, blinking its wall light in some code I did not know how to read. I did not call anyone. I took a photo of the trolley mechanism on my phone and described what the door was doing.

The first answer pointed me toward the safety sensors mounted near the bottom of the track, the small photo-eye pair that stops the door if anything crosses the beam. A second photo showed me exactly which part of the sensor housing to look at and where the status light actually sits. I nudged the sensor bracket until the blinking stopped. The door closed. Total time from first photo to working door: about two minutes.

What actually happened

Nothing about this required a frontier model. A model trained specifically on home repair and appliance diagnostics would likely have done the same job for a fraction of the compute cost. Voice input would have worked as well as photos and text. The interface was not the interesting part.

What mattered was a model that could look at a photo, recognize a mechanical assembly it had never seen labeled, connect it to a class of known failure modes, and tell me exactly where to look next. That is a narrow, well-defined reasoning task. It does not need the biggest model available. It needs a model with the right training data and the ability to reason across an image and a follow-up question in the same conversation.

The technology to close this gap already exists. A garage door proved it in two minutes. The question is whether appliance makers will build toward it.

What "AI Mode" actually means on my washing machine

Compare that two-minute diagnosis to what appliance makers are actually shipping. My LG washing machine has an "AI Mode." It is a default setting, not a capability I asked for or can meaningfully interact with.

  • It has never told me my hot water heater is showing early signs of failure.
  • It has never flagged a sensor or seal drifting out of spec before it caused a bigger problem.
  • It definitely has not told me my shirts are out of fashion, and I am only half joking that this would be a lower bar than catching a real mechanical fault early.

"AI Mode" has become a label on a spec sheet, not a description of what the device does. The gap is not technical. Small multimodal models capable of diagnosing a mechanical fault from a photo are already good enough for consumer hardware. The gap is that almost no appliance maker has connected a camera, a sensor feed, and a reasoning model into something that behaves like the two-minute garage door fix, running continuously, in the background, before something breaks instead of after.

The pattern enterprises should recognize

This is not a fringe capability anymore. Market research firm Market.us found that nearly 60 percent of enterprise applications built in 2026 already combine two or more data modalities, text, image, audio, or video, in a single model call. The garage door is the same shift showing up in a driveway instead of a data center.

Home security cameras already watch constantly. Google Nest already does pieces of predictive awareness today, flagging an alarm sound or an offline device. The next step is not a bigger model watching more footage. It is a home operating system where a sensor does not just observe, it diagnoses, the way my garage door camera moment did on request.

I have written before about mimik's push to run AI inference at the edge rather than routing everything to the cloud, and about how vision models finally crossed a usability threshold this year, in an earlier look at multimodal vision models. The garage door is the same trend showing up in a driveway instead of a data center. The infrastructure for ambient, diagnostic home AI is closer to ready than the appliances built on top of it.

This is the same distinction enterprise buyers keep missing when they evaluate AI vendors. A checkbox labeled "AI-powered" tells you nothing about whether the system reasons across modalities, catches a fault pattern early, or just runs a default preset. The garage door did not need a frontier model. It needed a builder who actually wired the reasoning to the sensor.

CIO/CTO Viability Question

When a vendor tells you their product is "AI-powered," ask them the same question my washing machine has never answered: what specific failure, three weeks before it happens, will your system actually catch and tell me about? If the answer is a feature roadmap instead of a working example, you are buying a label, not a capability.

Principal Research Director, Info-Tech Research Group · Former Adjunct Professor, Georgetown University, Entrepreneur in Residence, Stony Brook University, NY
Market.us. "Multi-Modal AI Platform Market Size." Market.us, 2026. market.us.
LG. "LG ThinQ AI Mode." LG Electronics, 2026. lg.com.
Google. "Google Nest." Google Store, 2026. store.google.com.
Bellamkonda, Shashi. "mimik Launches mimOE Studio." shashi.co, May 2026. shashi.co.
Bellamkonda, Shashi. "The AI That Finally Learned to See." shashi.co, Apr. 2026. shashi.co.
Disclaimer: This blog reflects my personal views only. Content does not represent the views of my employer, Info-Tech Research Group. AI tools may have been used for brevity, structure, or research support. Please independently verify any information before relying on it.