class

OpenAI::RealtimeSessionCreateResponseTurnDetection

Inherits JSON::Serializable < Reference < Object

Configuration for turn detection. Can be set to null to turn off. Server VAD means that the model will detect the start and end of speech based on audio volume and respond at the end of user speech.

Constructors

new(type_value : String | Nil = nil, threshold : Float64 | Nil = nil, prefix_padding_ms : Int64 | Nil = nil, silence_duration_ms : Int64 | Nil = nil)
Source
new(*, __pull_for_json_serializable pull : JSON::PullParser)
Source

Instance methods

prefix_padding_ms

Amount of audio to include before the VAD detected speech (in milliseconds). Defaults to 300ms.

Source
prefix_padding_ms=(prefix_padding_ms : Int64 | Nil)

Amount of audio to include before the VAD detected speech (in milliseconds). Defaults to 300ms.

Source
silence_duration_ms

Duration of silence to detect speech stop (in milliseconds). Defaults to 500ms. With shorter values the model will respond more quickly, but may jump in on short pauses from the user.

Source
silence_duration_ms=(silence_duration_ms : Int64 | Nil)

Duration of silence to detect speech stop (in milliseconds). Defaults to 500ms. With shorter values the model will respond more quickly, but may jump in on short pauses from the user.

Source
threshold

Activation threshold for VAD (0.0 to 1.0), this defaults to 0.5. A higher threshold will require louder audio to activate the model, and thus might perform better in noisy environments.

Source
threshold=(threshold : Float64 | Nil)

Activation threshold for VAD (0.0 to 1.0), this defaults to 0.5. A higher threshold will require louder audio to activate the model, and thus might perform better in noisy environments.

Source
type_value

Type of turn detection, only server_vad is currently supported.

Source
type_value=(type_value : String | Nil)

Type of turn detection, only server_vad is currently supported.

Source