The traditional modality used for Human-Computer Interaction (HCI) involves keyboard and mouse input, which limits the naturalness and accessibility of human interactions with computing systems. This paper describes the design and implementation of a multimodal intelligent assistant that allows hands-free communication with computers through voice interaction, hand gesture recognition, and artificial intelligence. The system consists of speech recognition, a Large Language Model (LLM) reasoning component for providing answers to natural-language queries, a Text-to-Speech (TTS) component for spoken answers, a voice dictation component for hands-free typing, screen awareness component with AI-vision capabilities for analyzing screen content, and a lightweight floating UI for showing the state of the assistant. Hand gesture recognition is implemented as an additional way to perform control actions along with natural-language voice commands. Multilingual switching between English and Telugu languages is implemented explicitly to avoid unreliable language detection. This paper presents the architecture and methodology of the proposed system and current status of its implementation, specifying implemented components, components in development, and proposed for future development. Instead of inventing performance metrics, The paper highlights qualitative findings based on tests carried out on the functioning prototype. The objective is to show how the combination of voice, gestures, and AI-based reasoning can bring human computer interaction to a more natural and inclusive model.
Keywords
Human Computer Interaction, Multimodal Interaction, Speech Recognition, Hand Gesture Recognition, Large Language Models, Desktop Automation, Text-to-Speech, Screen Understanding.
This research paper has discussed the design and current development of a multimodal intelligent assistant, which integrates the features of voice interaction, hand gesture recognition, AI/LLM reasoning, desktop automation, voice dictation, multilingual capability, and screen awareness into a single framework of Human-Computer Interaction. The currently developed modules, such as voice command processing, AI-based answering questions, Text-to-Speech, voice dictation, explicit multilingual capability, and floating status bar, have been discussed, together with those which are under development, including full screen awareness and complete hand gesture recognition. By taking the voice and gesture interaction as complementary modes of interaction and basing the responses of AI on the structured command-processing pipeline, the proposed system tries to reach a hands-free way of computer interaction.
References
[1]A. de Barcelos Silva, M. M. Gomes, C. A. da Costa, R. da Rosa Righi, J. L. V. Barbosa, G. Pessin, G. De Doncker, and G. Federizzi, "Intelligent personal assistants: A systematic literature review," Expert Systems with Applications, vol. 147, p. 113193, 2020, doi: 10.1016/j.eswa.2020.113193. [2]C. M. Rebman, M. W. Aiken, and C. G. Cegielski, "Speech recognition in the humanâcomputer interface," Information & Management, vol. 40, no. 6, pp. 509â519, 2003, doi: 10.1016/S0378-7206(02)00067-8. [3]F. F. Ambadar and J. J. L. Martinez, "Improving the performance of speech-gesture multimodal interface in non-ideal environments," Procedia Computer Science, vol. 216, pp. 587â596, 2023, doi: 10.1016/j.procs.2022.12.173. [4]Y. Deng, L. Fang, and X. Huang, "Research on multimodal human-robot interaction based on speech and gesture," Computers & Electrical Engineering, vol. 72, pp. 443â454, 2018, doi: 10.1016/j.compeleceng.2018.09.014. [5]A. Mahmood, J. Wang, B. Yao, D. Wang, and C.-M. Huang, "LLM-powered conversational voice assistants: Interaction patterns, opportunities, challenges, and design guidelines," arXiv preprint arXiv:2309.13879, 2023.
đ How to Cite This Paper
Y. Venkata Krishna, K. Chandu, K Rohith Kumar, G.Mohan Manjunadh Reddy, M. Manoj Reddy (2026). Voice and Hand Gesture-Based Multimodal Intelligent Assistant for Human-Computer Interaction. International Journal of Computer Science Engineering Techniques, 10(5), 266â271. ISSN: 2455-135X. DOI: https://doi.org/10.5281/zenodo.23281325