A Theory for a Multi Modal, Lightweight Personal AI that I don’t have the Skill or the Knowledge to build, Help Wanted

 

Background

I do not generally see myself as an AI user or necessarily as an advocate of the current implementations of AI systems, however I dream of a day where personal AI assistants exist that are light on hardware requirements, responsive to user interactions in a way that feels natural, but also still have the capacity to respond to complex queries and requests. The theory I include here has been logically checked with Google Search’s Gemini AI, so it might be incomplete or otherwise ineffective, but I have no way of checking or testing my theory myself and would highly appreciate some feedback from experienced AI developers as to how much potential there is in my theory.

I understand that my theory might not actually hold weight and that Gemini’s responses to my queries might be incorrect, or influenced by the positive reinforcement that can be built into AI models to encourage engagement.

The outlined theory is kept at a high level to compress the primary ideas and thought processes behind it.

The Theory

I recently found myself thinking about how I could develop such an AI system that could operate as described above, a local, personal AI that does not require every bit of data about me to be sent to a massive cloud service.

Considering some of the limitations of the current approaches to AI, I realised that current technologies require massive amounts of memory and processing capabilities, even when running a small local model. To think about a way to improve this situation, I considered how the human brain works. In the human brain, large amounts of input data are received, but only a very small proportion of this data is actively processed for conscious thought. The brain naturally filters this information down to the key components needed for processing, and all of this happens quickly.

Now most people naturally convey information through speech. Speech relies on a series of core sounds that are isolated from background noise by the brain, with each sound pattern relating to a learned context. Equally, these sounds can be altered based on the tone and emotion being conveyed by a speaker, further adjusting the context of the pattern when listened to.

So, on asking Gemini about the potential performance savings that could be made by using filtered audio input looking at the core components of speech (phonemes, pitch and context), and audio queues as the primary source of processing and organising data, Gemini responded that this could introduce significant performance gains over current approaches, and is actively being researched. However, the expectation of AI systems to be able to handle complex tasks still presented a problem, storage and processing requirements.

Currently, any AI system that is asked to perform a complex task requires a large amount of resources to undertake that task, preventing it from being small, and reducing the potential of it being able to function locally.

To get around this bottleneck, and thinking about a general user’s experience, most day to day tasks are normally quite simple. For example, setting calendar appointments, checking the news and weather, interacting with smart appliances. For these scenarios, a personal AI doesn’t need to know the information itself, or process massive amounts of information, it collects it from external resources, or relies on a small set of learned information. So how could it respond to complex queries?

What if the personal AI acted as an interface and a broker? It then wouldn’t need the resources to handle complex tasks locally as they would be handled externally, but for simple interactions, it would use a local resource base to make actions and respond to the user.

This approach would drastically reduce the local resource overheads required to run a personal AI model, reducing it to a smart chatbot, while still being able to take advantage of the more powerful back-end systems of current commercial AI models. It would also reduce the risks associated with privacy and security as most personal information is only ever processed locally.

To extend this, using both the audio based token approach, and the hybrid processing approach, it appears possible to be able to create a personal AI that can be responsive to normal interactions, while maintaining the ability to perform complex tasks by integrating with either current AI systems, or a bespoke service that acts similarly to current systems that are structured to cooperate with the personal AI’s.

On considering the types of user interactions with an AI system, a user will want to be able to interact in more ways than just speech, for example through text using a GUI, or through camera interactions. With the move to audio tokens, it appears as though audio tokens could be used for other input types as well, and possibly in more performant ways than current approaches, resulting in a shared token type that can be utilised by many different types of input. My theory on this would be to utilise software agents that translate an input type into the shared token type to allow different devices and input types to effectively utilise the same standard approach to processing information. The agents would not however perform any AI processing themselves, they would collect and pre-process the information into the shared token format, and relay it to a central personal AI host that would then process and return results through the appropriate vectors, reducing the resource requirements of the software agent hosts. This would be handled through current network technologies to allow for local network communications between locally installed agents, or through a VPN, or similar technical approach for mobile devices away from the local network.

This approach allows for fluid multi-modal interactions to occur between the user, the software agents, and the local personal AI hub, meaning that regardless of the input type, the personal AI can maintain the context of a user’s interactions with it. This would also improve the user experience by maintaining the target of the interaction as the personal AI.

So the result is a system for a personal AI that should be able to:

  • operate locally for daily tasks,

  • maintain a relatively small local resource footprint for use on current domestic hardware,

  • maintain the capability to respond to complex queries through using current AI services, or similar,

  • representatively respond naturally to user interactions with less latency,

  • be able to receive inputs from multiple sources and types of data source,

  • protect user privacy and security through keeping personal information on the local device,

  • operate in a more performant way when compared to current approaches for everyday tasks,

  • reduce the power and resource requirements of most interactions, and the heavy use of current models for simple interactions.

While my theory is likely incomplete, or founded on erroneous responses from Gemini, my requests to Gemini were specific, raised as hypothetical, and requested possible problems with my queries. Where a query resulted in a negative or incomplete response, I either refined my idea or request to mitigate potential problems with the theory.

Through the conversation, I received several outputs from Gemini to try and record the thought process and results. I do not know how beneficial this idea or approach will be, or what any actual performance gains might be in experimenting with this approach which is why I am making this post here.

If anyone would like to learn more about my idea, or the outputs I received in “testing” the theory against Gemini, let me know. I would like to see something like this be created, and to have a part in creating it if possible.

If my theory has no potential, equally let me know why you think this wouldn’t work, perhaps there’s more to think about.

Email me at: crazyhoundgamedesign@gmail.com

Or message me on X: https://x.com/CrazyHoundGames


Comments

Popular posts from this blog

Indie Dev Interview: Alnutmob Studios

Blog Schedule September - November 1st

Upgrading My Art #2 - Practice Makes Perfect, but Style?