KAIST Develops Smartphone AI That Uses Past Solutions to Tackle Similar Problems
en-GBde-DEes-ESfr-FR

KAIST Develops Smartphone AI That Uses Past Solutions to Tackle Similar Problems


A new AI technology has been developed that retains and reuses knowledge, much like writing down a solution in a notebook and applying it to similar problems rather than asking an expert for help each time. Researchers at KAIST have developed a way for small AI models running on smartphones to store and reuse knowledge from a large server model. The approach reduced server calls by an average of 55.61% compared with a non-cumulative approach while maintaining high accuracy, suggesting that mobile AI could make faster decisions with less server support as it encounters similar problems.

KAIST (President Choongsik Bae) announced on September 21 that a research team led by Professor Jae-Gil Lee from the School of Computing has developed CURE (Cumulative Knowledge Reuse), a technology that enables a small AI model on a device to work efficiently with a large model on a server.

Smartphones have limited processing power and memory, so they typically use small, lightweight AI models. These models can handle simple tasks quickly but may be less accurate when identifying complex or unfamiliar images.

Sending every input to a powerful server model can improve accuracy, but transferring data and waiting for a response takes time. It also adds to network traffic and the server’s computational workload, increasing delays for users and operating costs for service providers.

Researchers have explored a collaborative approach in which the on-device model assesses each input first and sends only difficult cases to the server. Without a way to retain the server’s knowledge, however, the device uses each answer once and may need the same help when a similar input appears. This is much like a student who fails to write down the solution to a difficult problem and has to ask for help again when faced with another of the same kind. The team set out to turn these one-time answers into knowledge the device could continue to use.

CURE first checks whether the on-device model can handle an input reliably. If it cannot, the system consults knowledge previously obtained from the server and stored on the device. It contacts the server only if that knowledge is also insufficient. The process follows three steps, moving from solving a problem independently to consulting previous lessons and, when necessary, asking an expert.

For example, an on-device model that cannot identify a car model in a photo may turn to the server for help. CURE uses the server’s prediction and the image’s features to update the device’s knowledge store. When the device later receives a photo of a similar car model, it can use that knowledge to identify it locally.

CURE does not keep a collection of the original photos processed by the server. Instead, it stores a summary of their shared features and differences. This is like noting a car’s distinguishing features, such as its body shape or headlight design, rather than memorizing the entire photo. The stored knowledge can therefore help the system recognize not only images it has already seen but also similar images it encounters for the first time.

The team tested CURE using vision-language models, which connect images with text to understand visual information. These models link what they see to language, much as people do.

The researchers used MobileCLIP2 on the device and EVA-CLIP, which has 18 billion parameters, on the server. Parameters are numerical values within a model that encode information learned during training. They help the model distinguish objects and identify relationships between them.

In tests on a range of image classification datasets, CURE achieved accuracy close to that of an approach that sends every input to the large server model. It also made an average of 55.61 percent fewer server calls than a device-server collaboration baseline that does not retain knowledge from previous server responses.

In end-to-end tests that included communication time, CURE ran up to 2.80 times as fast as the non-cumulative device-server collaboration baseline and up to 3.67 times as fast as the approach that sends every input to the server. Making fewer server calls reduced the time spent transferring data and waiting for responses.

CURE requires no retraining of either the on-device model or the server model. It leaves both models unchanged and adds a separate store for knowledge obtained from the server, making it adaptable to a range of AI models and services.

The technology could reduce repeated data transfers and server computation image recognition. This could mean shorter waits for users and lower operating costs for service providers.

Potential applications also include robots and wearables that need to recognize their surroundings with limited computing resources. Robots operating over slow or unreliable networks, for example, could use previously acquired knowledge to make more decisions locally without waiting for a server response.

The benefits in practice will depend on the device’s processing power and storage capacity, network conditions, and the characteristics of the input data. Further testing across different devices and environments is therefore needed.

“CURE allows a small on-device AI model to remember and reuse knowledge it has already obtained, rather than repeatedly asking the server the same question,” said Professor Jae-Gil Lee. “By maintaining high accuracy while reducing communication demands and response times, we expect it to help smartphones, robots, and wearables use powerful AI models more efficiently.”

Dr. Youngjun Lee, a postdoctoral researcher from the KAIST Institute of Information Electronics, was the study’s first author, and Professor Jae-Gil Lee from the School of Computing was the corresponding author. Co-authors were Doyoung Kim from Amazon, Junhyeok Kang from LG AI Research, and Professor Hwanjun Song from the KAIST Department of Industrial and Systems Engineering. The findings were presented on September 10 at the European Conference on Computer Vision (ECCV 2026), a leading international computer vision conference held in Malmö, Sweden, from September 8 to 12.

Paper title “CURE: Cumulative Knowledge Reuse for Efficient Device-Server Hybrid Inference in Vision-Language Models”, DOI: https://doi.org/10.1007/978-3-032-37627-5_21
Author information — Dr. Youngjun Lee (postdoctoral researcher, KAIST Institute of Information Electronics, first author), Doyoung Kim (KAIST graduate, now at Amazon, co-author), Junhyeok Kang (KAIST graduate, now at LG AI Research, co-author), Professor Hwanjun Song (KAIST Department of Industrial and Systems Engineering, co-author), Professor Jae-Gil Lee (KAIST School of Computing, corresponding author)

This work was supported by Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2020-II200862, DB4DL: High-Usability and Performance In-Memory Distributed DBMS for Deep Learning, 50% and No. RS-2022-II220157, Robust, Fair, Extensible Data-Centric Continual Learning, 40%), the InnoCORE pro-gram of the Ministry of Science and ICT (AI Meta-Scientist, N10260110, 10%), and Samsung Electronics Co.,Ltd.(IO251216-14634-01).
Paper title “CURE: Cumulative Knowledge Reuse for Efficient Device-Server Hybrid Inference in Vision-Language Models”, DOI: https://doi.org/10.1007/978-3-032-37627-5_21
Author information — Dr. Youngjun Lee (postdoctoral researcher, KAIST Institute of Information Electronics, first author), Doyoung Kim (KAIST graduate, now at Amazon, co-author), Junhyeok Kang (KAIST graduate, now at LG AI Research, co-author), Professor Hwanjun Song (KAIST Department of Industrial and Systems Engineering, co-author), Professor Jae-Gil Lee (KAIST School of Computing, corresponding author)
Attached files
  • Figure 1. Comparison between a non-cumulative hybrid approach and CURE
  • Schematic illustration of the study (AI-generated)
Regions: Asia, South Korea, Europe, Sweden, North America, United States
Keywords: Applied science, Artificial Intelligence, Computing, Engineering, Technology

Disclaimer: AlphaGalileo is not responsible for the accuracy of content posted to AlphaGalileo by contributing institutions or for the use of any information through the AlphaGalileo system.

Testimonials

For well over a decade, in my capacity as a researcher, broadcaster, and producer, I have relied heavily on Alphagalileo.
All of my work trips have been planned around stories that I've found on this site.
The under embargo section allows us to plan ahead and the news releases enable us to find key experts.
Going through the tailored daily updates is the best way to start the day. It's such a critical service for me and many of my colleagues.
Koula Bouloukos, Senior manager, Editorial & Production Underknown
We have used AlphaGalileo since its foundation but frankly we need it more than ever now to ensure our research news is heard across Europe, Asia and North America. As one of the UK’s leading research universities we want to continue to work with other outstanding researchers in Europe. AlphaGalileo helps us to continue to bring our research story to them and the rest of the world.
Peter Dunn, Director of Press and Media Relations at the University of Warwick
AlphaGalileo has helped us more than double our reach at SciDev.Net. The service has enabled our journalists around the world to reach the mainstream media with articles about the impact of science on people in low- and middle-income countries, leading to big increases in the number of SciDev.Net articles that have been republished.
Ben Deighton, SciDevNet

We Work Closely With...


  • The Research Council of Norway
  • SciDevNet
  • Swiss National Science Foundation
  • iesResearch
Copyright 2026 by AlphaGalileo Terms Of Use Privacy Statement