{
  "contents": [
    {
      "role": "user",
      "parts": [
        {
          "text": "Two complete sung advertorial story-songs about the same story: Attachment 1 is in English, Attachment 2 is a Ukrainian adaptation of the same story with a different product. Listen to BOTH whole songs end to end. Task: a songcraft and vocal-delivery comparison at the META level. Specifically analyze, with at least 10 timed landmarks per song: (a) how each singer ATTACKS and RELEASES words (crisp consonant endings vs open sustained vowels; which words get lengthened and whether they are the important words or unimportant ones); (b) syllable-length shaping: where syllables are clipped, where stretched, and whether stretch coincides with meaning; (c) melodic play with words: moments where the melody makes a joke, a turn, a question, a surprise, or lets a word \"land\" (describe the gesture); (d) rhythm: whether vocal phrases sit on the beat grid or float speech-like across it, and where syncopation or off-grid placement creates emphasis; (e) phrase endings and breath: what happens musically between phrases (band answers, held chord, silence, immediate next line); (f) how the arrangement reacts to story turns (drops, stops, fills, register changes); (g) register and timbre choices per scene; (h) form and pacing of information: where each song is fast and where it slows down, and whether density tracks drama. Then give: transferablePrinciples[{principle,evidenceAttachment1,evidenceAttachment2,howToApplyInUkrainianWithoutFragmentingSentences}], languageInherentDifferences[] (things caused by English vs Ukrainian word structure, not by the singer), strongestGesturesEach[], weakestGesturesEach[], whatAttachment2AlreadyDoesAsWellOrBetter[]. Constraint for recommendations: Ukrainian must stay natural, fully sung and intelligible; do NOT recommend chopping sentences into telegraphic fragments or inserting long instrumental gaps after every line. Return ONLY JSON (no prose outside JSON), <=1900 words, English. Required keys: mediaAccess(boolean), audioAccess(boolean), attachmentsHeard[{attachment,firstWordsHeard,lastWordsHeard,durationEstimateSeconds}], then the analysis keys listed below. No script is supplied: quote only words you actually hear and mark uncertain words with (?). Use attachment number plus LOCAL seconds of that attachment. Do not rate quality with numbers, do not assume either language or the original is better, do not reward louder mastering, more notes or softer timbre. Both are real songs: identify concrete transferable mechanisms, not taste."
        },
        {
          "text": "Attachment 1: Attachment 1; complete English song, 194.5s, 56kbps mono listening proxy with fixed gain (do not grade codec quality); local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 2: Attachment 2; complete Ukrainian song, 203.2s, 56kbps mono listening proxy with fixed gain (do not grade codec quality); local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        }
      ]
    }
  ],
  "stream": true,
  "generationConfig": {
    "thinkingConfig": {
      "includeThoughts": false,
      "thinkingLevel": "high"
    }
  }
}
