
Closed
Posted
Paid on delivery
I need a proof-of-concept mobile web application that works smoothly on both Android and iOS. The core idea is simple but technically rich: a visitor points their phone at a public monument, the camera captures an image, the app recognises the figures depicted, and then each of those historical characters speaks back in a distinct, period-appropriate voice. Here is the flow I have in mind: 1. Vision: Use a lightweight on-device or edge computer-vision model to identify every figure in the photo. Accuracy, speed, and the ability to handle outdoor lighting are crucial. No text recognition is required—just figure identification. 2. Voice design: For each identified figure, generate an era-authentic voice with ElevenLabs’ Voice Design API. Think marble statues suddenly speaking in first-person tones that match their century. 3. Multi-voice orchestration: Feed the voices into an ElevenLabs Agent so multiple characters can converse naturally with the user and with one another. The agent needs to manage context, respond in real time, and switch seamlessly between speakers. 4. Real-time dialogue: Stream audio over WebRTC so latency stays low enough that the exchange feels like a live conversation. When the user speaks, transcribe locally (or with Whisper) and pass the text back to the agent for a response. 5. Optional research depth: If you can layer in quick web look-ups to enrich the agent’s knowledge about lesser-known statues, that would be a plus, but keep the main loop fast. Acceptance criteria • A single mobile-friendly web URL that I can open on both Android and iOS browsers. • Take/choose a photo of a statue, identify figures, and show their names on screen. • Tap “Talk” and hold a natural voice conversation where each figure replies in its own, era-authentic voice with round-trip latency under two seconds. • Source code, environment setup notes, and a short README that explains how to swap out models or keys. If you have prior experience with ElevenLabs, WebRTC, or on-device vision, let me know—that will speed up the build. We have the directions how to make the app
Project ID: 40628726
170 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
170 freelancers are bidding on average $3,935 USD for this job

Hi, I saw your JD and having 12+ years of experience in AI-powered web applications, computer vision, real-time communication, and LLM integrations. I understand you need a mobile-first web application that identifies historical figures from monument photos, generates authentic voices using ElevenLabs, enables real-time multi-character conversations via WebRTC, and delivers a seamless experience on both Android and iOS browsers with low latency. Key Features: * Mobile Web App * Computer Vision * Figure Detection * ElevenLabs API * Voice Design * Multi-Agent Chat * WebRTC Streaming * Whisper STT * Real-Time Audio * Context Memory * Admin Panel * Secure Hosting I have experience building AI applications with LLMs, computer vision, speech-to-text, text-to-speech, WebRTC, and real-time AI workflows. Based on your provided architecture and implementation guide, I can build a scalable proof of concept with clean, well-documented code, environment setup, and an extensible architecture for future enhancements. Let’s Chat… Thanks
$3,200 USD in 24 days
9.4
9.4

Hello, I'm excited by this innovative proof-of-concept for a mobile web app that brings history to life. I understand you're looking for a solution that accurately identifies figures on monuments using on-device vision, generates period-appropriate voices via ElevenLabs, and orchestrates real-time, low-latency conversations with users over WebRTC. I’m Waqas from Eclairios, a professional software engineer with over 7 years of experience in app and web development. I have successfully completed 128 projects, earning a 5.0 rating from satisfied clients. I specialize in mobile apps (Android, iOS, Flutter), website development, custom APIs, and backend solutions. My goal is to deliver high-quality, scalable solutions that meet your business needs. Why hire me? ★ 100+ Projects Completed with 5-star rating. ★ 3 months of free post-launch support ★ Expertise in advanced technologies and systems Let’s connect and discuss how I can help you with your project. Best regards, Waqas
$3,420 USD in 7 days
8.4
8.4

Hi — Elias here from Miami. I see you're looking to create a proof-of-concept mobile web application for interactive conversations with a statue. The goal is to deliver an engaging user experience on both Android and iOS. What usually matters most here is ensuring seamless functionality across platforms while managing complex interactions. A common issue in systems like this is integrating real-time communication and image recognition, which can introduce performance and scalability challenges. My approach would involve structuring a robust backend that supports real-time communication via WebRTC, ensuring smooth interactions. I'll focus on maintainability by using modular components, making future updates easier. From my experience with similar interactive applications, I understand the importance of creating an intuitive user interface that enhances engagement. A few questions to better understand the scope: Q1 – What specific user roles and permissions will the app require? Q2 – Are there particular backend integrations you envision for data handling? Q3 – How do you plan to manage user authentication and privacy? Happy to discuss the details and suggest the best technical approach. Looking forward to hearing from you.
$4,500 USD in 14 days
8.2
8.2

Hello, Greetings Hope you are doing well. I am a Mobile Application developer with over 5 years of experience. I am confident in my ability and skills to develop high-quality Mobile apps and would like to work on your statue conversation app project. I will complete the work as per your requirements. I want to discuss more this project to prepare the final concept. So let’s discuss this in detail over chat then will make plans to start work on it. Waiting for your earliest reply. Thanks. Shubham
$3,000 USD in 35 days
8.0
8.0

Hi there, This is a strong proof-of-concept idea, and the main challenge is keeping the experience fast and reliable on real mobile browsers rather than overbuilding the first version. I can help build it as a mobile-friendly web app with a clean React/Node.js architecture, so Android and iOS users can open one URL, capture or upload a monument photo, see identified figures, and start a low-latency voice conversation. For the POC, I would keep the vision pipeline practical: use a lightweight model or edge API focused on known statue/figure recognition, with clear fallback handling for poor lighting or uncertain matches. On the voice side, I would integrate ElevenLabs Voice Design and Agent workflows so each recognized character can have a distinct voice while the app manages speaker switching and conversation context. For real-time interaction, I would use WebRTC audio streaming where supported, with Whisper or browser-compatible transcription depending on latency and device constraints. I would also structure the code so models, API keys, and voice mappings can be swapped without rewriting the app. I have extensive mobile and full-stack experience across Flutter, native iOS/Android, React, and Node.js, which is useful here because the biggest risks are mobile browser behavior, audio permissions, latency, and maintainability. I’d be happy to discuss your existing directions and turn them into a focused, testable POC.
$4,000 USD in 35 days
7.9
7.9

As an AI & Embedded Systems Engineer with a mastery of both Android and iOS platforms, I offer precisely the skillset you're seeking for your Interactive Statue Conversation App. My understanding of computer vision and image recognition allows me to tackle the challenge of accurately identifying the figures on public monuments swiftly and regardless of outdoor lighting conditions. Moreover, my prior work in Voice Design adds an era-appropriate voice component that I believe will add a unique level of immersion to your app. I'm genuinely excited about harnessing these capabilities for multiple characters to converse with gesture onlookers convincingly. Having worked extensively with ElevenLabs, I bring in familiarity with their Voice Design API, enabling seamless integration of authentic voices and maintaining low-latency streaming during real-time dialogues via WebRTC for an immersive user experience. Most importantly, as an end-to-end solution-driven engineer, I understand the significance of delivering a product with clear instructions for seamless swapping of models or keys. Rest assured, you won't just receive a fully-functional mobile web app, but also comprehensive handover documentation. Let's bring history alive through statues and create a proof-of-concept that truly captures imaginations - together!
$5,000 USD in 60 days
7.0
7.0

Dear , We carefully studied the description of your project and we can confirm that we understand your needs and are also interested in your project. Our team has the necessary resources to start your project as soon as possible and complete it in a very short time. We are 25 years in this business and our technical specialists have strong experience in Mobile App Development, iPhone, Android, Objective C, Computer Vision, WebRTC, ElevenLabs, Image Recognition and other technologies relevant to your project. Please, review our profile https://www.freelancer.com/u/tangramua where you can find detailed information about our company, our portfolio, and the client's recent reviews. Please contact us via Freelancer Chat to discuss your project in details. Best regards, Sales department Tangram Canada Inc.
$4,160 USD in 5 days
7.2
7.2

Hi, I've built multi-modal mobile apps that blend computer vision with real-time voice—this is exactly our wheelhouse. Your statue recognition + ElevenLabs agent orchestration with WebRTC streaming is doable as a tight PoC in this budget window. Message me to discuss the vision model choice and voice latency targets. Regards, Nurul Hasan
$3,800 USD in 30 days
6.7
6.7

Hi, I can develop this proof-of-concept mobile application that works seamlessly on both Android and iOS browsers. The application will capture or upload a monument photo, identify the historical figures using an optimized computer vision model, and enable natural multi-character conversations with distinct era-appropriate voices powered by ElevenLabs. The solution will include low-latency WebRTC audio streaming, real-time speech transcription, and a scalable architecture so models or APIs can be easily replaced in the future. I have experience building AI-powered MVPs involving computer vision, real-time communication, and API integrations. Based on the provided direction, I can deliver the complete POC with clean source code, setup documentation, and deployment guidance. Let's discuss further. Regards, SNR
$3,000 USD in 10 days
6.9
6.9

Hi I have read your requirements and I am sure I will be able to help you. Please message me so that we will have detail technical discussion. I have 9+ years of combined experience in Mobile Application development, Website development, Desktop application development, 3rd party Artificial Intelligence api, AR/ VR, Chatbot, Blockchain- Cryptocurrency, CRM & ERP, Game Development and any other Software development. I am having expertise in Native on Android Java, kotlin and IOS Swift, and For Hybrid Cross platform on Flutter Dart & React- Native, and for web and backend on react js and node js, Python Django. Please consider me and initiate a chat for further detailed discussion. Regards, Anju
$3,000 USD in 45 days
6.6
6.6

Hi, This is an exciting AI application because it's much more than image recognition. The experience depends on combining fast vision inference, contextual LLM conversations, and real-time voice synthesis into one natural interaction. I would build the POC as a mobile-first web application using WebRTC for streaming audio, a lightweight vision model for statue recognition, Whisper (or equivalent) for speech-to-text, and ElevenLabs Voice Design and Conversational AI for multi-character dialogue. Each figure would maintain its own persona and conversation context while allowing seamless turn-taking between characters and the visitor. The solution will be modular, making it easy to swap vision models, LLMs, or voice providers as the project evolves. You'll receive the complete source code, deployment guide, environment configuration, and documentation for replacing models or API keys. I'm ready to begin immediately and can deliver the project in milestone-based phases with regular demos.
$4,500 USD in 45 days
6.0
6.0

Hello! As per your project description, you are looking to build an Interactive Statue Conversation App where visitors can point their phone at a public monument, have the system identify the historical figures depicted, and then interact with those characters through distinct period appropriate voices. The experience will combine computer vision, AI voice generation, multi character orchestration, speech recognition, and low latency real time audio into one engaging interactive experience that works smoothly across Android and iOS. My focus will be on keeping the core interaction fast and natural. The vision layer can identify the relevant figures from the camera image using an appropriate lightweight model, after which ElevenLabs Voice Design can generate character specific voices. The identified characters can then be orchestrated through an AI agent capable of maintaining conversation context, switching between speakers. I specialize in AI and ML integrations, computer vision, mobile web applications, LLM and conversational AI, ElevenLabs integrations, speech to text, text to speech, WebRTC, real time streaming, API development, and responsive cross platform experiences. Let’s connect to discuss the initial monuments, recognition approach, ElevenLabs workflow, real time conversation requirements, and POC expectations so we can create an immersive experience where historical figures feel genuinely present and conversational. Best regards, Nikita Gupta
$3,000 USD in 45 days
5.7
5.7

hello there , I can developed your mobile web app with the complete flow you described: monument image recognition ,historical figure identification , AI voice generation , real-time multi-character conversation. I have experience working with AI integrations, computer vision workflows, real-time communication, and API-based app. I understand the importance of keeping latency low while maintaining a natural experience. I can also help evaluate the best architecture for balancing on-device processing and cloud AI services. I would be happy to discuss the technical stack and start building the POC. thanks
$3,000 USD in 25 days
5.2
5.2

As an extensive Full-Stack Mobile and Web Developer, I possess the skills necessary to bring your Interactive Statue Conversation App to life. My proficiency in building end-to-end solutions makes me capable of managing every aspect of this project, from the frontend interface, using tools such as Flutter and React Native, to backend processing with Node.js and .NET technologies. Your preference for clean, scalable code aligns with my approach, ensuring a robust and maintainable deliverable. With multiple successful app developments under my belt, including those involving real-time engagement and integration with third-party APIs like ElevenLabs, I possess a comprehensive understanding of what it takes to meet and exceed project specifications. Rest assured that choosing me means prioritizing consistent on-time delivery without compromising quality or post-project support. In conclusion, let's leverage my experience, ownership mentality, and intrinsic motivation for excellence to bring your unique statue conversation concept to life attractively on both Android and iOS platforms.
$4,500 USD in 60 days
5.0
5.0

Hi, Aashiq here from Cape Town, South Africa. This project instantly caught my eye, so I had to reach out. I see you’re looking for a mobile web app that uses computer vision to identify statues and engages users with era-authentic voices. That’s a fascinating concept! I’ve helped businesses create engaging mobile experiences, leveraging technologies like WebRTC and ElevenLabs. My background in developing interactive applications ensures I can deliver the seamless, real-time conversation flow you envision. Feel free to ask for samples of similar projects I’ve completed. Based on what you mentioned, here is how we would approach the project: - Develop a lightweight on-device vision model for accurate figure identification. - Integrate ElevenLabs’ Voice Design API for authentic voice generation. - Use WebRTC for real-time audio streaming and dialogue management. Rest assured, I prioritize clear communication and will deliver a user-focused solution optimized for performance. Best Regards, Aashiq
$4,500 USD in 7 days
4.6
4.6

I believe I can provide you with the perfect solution for your interactive statue conversation app. With over 5 years of experience in both web and mobile app development, I have the skills necessary to create a functional and user-friendly application that smoothly runs on Android and iOS devices. In terms of your project's core idea, I have expertise in not only frontend and backend development but also other areas like API integration, database management, performance optimization, and UI/UX-friendly development. This means I can ensure that your application captures the image accurately even in outdoor lighting conditions while identifying the figures depicted. Regarding ElevenLabs, WebRTC, or on-device vision, though I'm yet to work directly with them, my versatility as a developer has enabled me to quickly learn and adapt to new technologies. As such, I'm confident that I can utilize my prior experience with similar tools and frameworks effectively to fulfill your project requirements. Let us collaborate and transform this intriguing concept into a reality!
$4,000 USD in 7 days
4.4
4.4

I can help you turn this into a strong proof of concept that feels polished rather than just experimental. The combination of statue recognition, multi-character voice replies, and real-time audio flow is very doable, but it needs careful orchestration to keep latency low and the experience smooth on mobile browsers. I’d keep the build lean around the core loop first: capture/select image, identify figures, display names, then launch a natural conversation flow with distinct voices. If useful, I can also structure it so the research-enrichment layer stays optional and does not slow down the main interaction.
$3,900 USD in 7 days
3.6
3.6

The core idea behind your project is not just technically rich but also unique. It requires an expert who can handle lightweight on-device or edge computer-vision models for outdoor lighting conditions, points where I thrive. The use of the Voice Design API from ElevenLabs is another area I'm well versed at, capable of providing era-authentic voice for each identified figure within your desired real-time dialogue range. Additionally, while I may not have prior experience with ElevenLabs or WebRTC, my ability to learn quickly and adapt to new technologies will ensure a smooth and timely delivery. You require someone who can work autonomously without compromising on regular updates and communication – this is what you get when you choose to work with me. Together, we'll add value to your project through clean, maintainable code for easy scaling and long-term success. Let's create an amazing piece of technology that adds depth to historical conservation.
$4,500 USD in 10 days
3.5
3.5

Hello, Ricardo from Buenos Aires here. I've handled similar challenges before. "INTERACTIVE STATUE APP" - you need an AI experience that turns historical figures into natural conversations. Your goal is to create a mobile-friendly app where users can recognize statues and have realistic voice conversations with historical characters. I understand the need for computer vision, AI voice interaction, real-time responses, and a smooth experience on both Android and iOS. My creative idea is to combine image recognition with intelligent voice agents and natural dialogue flow to make the statues feel alive. I can help build a strong proof-of-concept with a clear structure and engaging user experience. As we move forward, let us remember that great AI applications should combine advanced technology with simple and enjoyable user interaction. Feel free to message me if you would like to discuss the details. I’d be happy to review your requirements and suggest the best next steps.
$3,400 USD in 20 days
3.5
3.5

Leveraging the rich experience of our exceptional development team at Web Crest, we present ourselves as the best choice for building your imaginative "Interactive Statue Conversation App". With our core expertise lying in Android and mobile app development, mixed with proficiency in edge-computer vision and AI, we possess all the technical prowess to make this application a reality. Your project's unique requirement for figure identification, era-authentic voice generation, multi-voice orchestration, and real-time dialogue align closely with our skill set. We have successfully executed projects merging similar cutting-edge technologies like ElevenLabs’ Voice Design API and WebRTC. Our proficiency in React Native and Flutter ensures a seamless performance on diverse platforms. Moreover, we don't just stop at project completion; we offer long-term technical support to ensure that even after deployment, your app functions smoothly. So why settle for less when you can choose a team that not only knows how to create an impressive mobile web application but also caters to all your post-development needs? Join hands with Web Crest and let us turn your extraordinary concept into an enduring reality!
$3,000 USD in 7 days
2.8
2.8

copenhagen, Denmark
Payment method verified
Member since Apr 21, 2016
$30-250 USD
£1500-3000 GBP
$250-750 USD
$30-250 USD
$30-250 USD
$30-250 USD
$10-30 USD
$250-750 USD
₹250000-500000 INR
₹5000-5001 INR
$750-1500 USD
₹150000-250000 INR
$250-750 USD
$1500 HKD
€30-250 EUR
£750-1500 GBP
$3000-5000 USD
$30-250 USD
₹12500-37500 INR
$600-1200 USD
₹100-400 INR / hour
$50-90 USD
$750-1500 USD
₹12500-37500 INR
$1500-3000 USD