Testing two-way audio on robots: echo, noise and operating conditions

Map the robot audio path, test echo and overlapping speech, and define RTC integration responsibilities using the conditions of the finished device.

A robot that can make a call while sitting on a desk may behave differently while moving or performing its normal work. Speakers, microphones, fans, motors and the enclosure all form part of the audio environment. RTC evaluation should therefore include the complete robot, rather than relying on a recorded speech sample or an idle development board.

Map every audio processing stage

Draw the path from microphone capture through device processing, RTC input, transmission and remote playback. Draw the return path separately. Note the sample format, channel count and the component responsible for echo cancellation, noise reduction and volume control.

If hardware, the operating system and the SDK all offer audio processing, confirm the intended configuration with the relevant teams. Turning every option on makes it harder to understand the result. When supplying external audio to an SDK, verify the available interface and how the relevant echo-processing component receives any playback reference it requires.

This is also a useful place to document ownership. Device engineers should be able to identify where local audio can be inspected. The RTC integration team should identify what enters and leaves the communication layer. Without that distinction, an acoustic issue can easily become a long-running network investigation.

Test echo and overlapping speech separately

Start with the remote participant speaking alone. Listen for that person’s voice returning through the robot’s speaker and microphone path. Next, have a person near the robot speak while the remote participant remains quiet. Check intelligibility and practical speaking volume.

Finally, let both people speak at once. Look for speech being heavily reduced, cut off or delayed until the other person stops. An enabled echo-cancellation setting is not an acceptance result. The media-capture specification describes capabilities and constraints; the finished device still needs an actual conversation test.

Include normal robot activity

Cover stationary operation, movement, turning, fan activity and representative processing load. Use the intended speaking distance and orientation. Record volume, power mode and enclosure version so another engineer can reproduce the conditions. If different microphones or headsets are supported, test switching between them.

Include interrupted workflows too: a call arriving while a device prompt is playing, the application changing state, audio reopening after a network interruption, and a new call immediately after the previous one ends. Confirm that microphone and speaker use return to the intended state after each call.

Agree on responsibilities and evidence

The device team supplies the audio interfaces and acoustic conditions. The RTC team confirms communication interfaces and diagnostic data. Product owners define the conversation that must work and the acceptance criteria. Classify failures as local capture, local playback or transport before assigning corrective work.

These are general testing recommendations. They do not promise a particular robot’s audio quality. For an RTC integration and pricing assessment, provide the system, chipset, audio path and typical operating environment.

Reference: W3C Media Capture and Streams describes browser media capabilities. Native device interfaces must be verified against the selected SDK.

← Back to insights

LET’S CONNECT

Start with your device and use case.

Tell us your device type, communication needs and project stage to discuss RTC integration and pricing.

Discuss integration & pricing ↗