CertKeen
Microsoft AzureBeta · expanding bank

Azure AI Fundamentals (AI-901) Practice Exam

Practice questions for the Microsoft Azure AI Fundamentals (AI-901) certification: responsible AI principles (fairness, reliability and safety, privacy and security, inclusiveness, transparency, accountability); how generative models work, choosing models by capability, deployment types and parameters such as temperature, top-p and max tokens; AI workloads including generative and agentic AI, text analysis, speech, computer vision, image generation and information extraction; and building solutions with Microsoft Foundry, including system and user prompts, the Foundry portal playgrounds, chat clients with the Foundry SDK and Responses API, prompt agents with tools and conversations, Azure Language and Azure Speech in Foundry Tools, multimodal audio and image input, image generation, and document, image, audio and video extraction with Azure Content Understanding. Every question includes a written explanation.

100 questions · 12 free preview

$19 · lifetime access
Try free sample

Studying more than one? All Microsoft Azure exams for $29 · every exam for $79

Free sample questions

  1. Sample · question 1 · Entity linking to disambiguate mentions

    An encyclopedia app built by Larkhill Media sees the word "Mercury" in articles and must decide whether each mention means the planet, the element or the Roman god, and then link to a reference page about the right one. Which Azure Language feature does this?

    • A.Language detection
    • B.Entity linkingcorrect
    • C.Sentiment analysis
    • D.Key phrase extraction

    Why: Entity linking disambiguates the identity of an entity found in text and returns a link to a knowledge base entry, such as a Wikipedia article, for the intended meaning. Key phrase extraction lists main topics without resolving what they refer to, language detection identifies the language, and sentiment analysis measures opinion.

    Open this question on its own page →
  2. Sample · question 2 · Pronunciation assessment for learners

    A language-learning startup named Tongueworks wants students to read sentences aloud and receive feedback on how accurately and fluently they pronounced each word. Which Azure Speech capability is designed for this?

    • A.Pronunciation assessmentcorrect
    • B.Speech synthesis with SSML
    • C.Speaker diarization
    • D.Batch transcription

    Why: Pronunciation assessment evaluates spoken audio and gives feedback on accuracy and fluency, which suits language learners. Speech synthesis produces audio from text, diarization separates who spoke when, and batch transcription converts large volumes of recordings to text without scoring pronunciation.

    Open this question on its own page →
  3. Sample · question 3 · Spoken language identification

    A hotline run by Brookvale Council receives calls in English, Polish and Urdu. Before transcribing each call, the system must determine which of these languages is being spoken. Which Azure Speech feature should it use?

    • A.Opinion mining
    • B.Custom neural voice
    • C.Key phrase extraction
    • D.Language identificationcorrect

    Why: Language identification compares spoken audio with a list of candidate languages to determine which one is being spoken, and it can be combined with speech-to-text or speech translation. A custom neural voice is for synthesis, and key phrase extraction and opinion mining analyze text rather than identifying the language of audio.

    Open this question on its own page →
  4. Sample · question 4 · Text-to-speech avatar videos

    Kittering Pharma wants short training videos in which a photorealistic digital presenter speaks a script with a natural voice, without filming a real person. Which Azure Speech capability supports this?

    • A.Text-to-speech avatarcorrect
    • B.Speech translation
    • C.Optical character recognition
    • D.Custom speech model training

    Why: Text-to-speech avatar converts text into a video of a photorealistic human speaking with a synthesized voice, either in real time or asynchronously. Speech translation converts speech between languages, custom speech improves recognition accuracy, and OCR reads text from images.

    Open this question on its own page →
  5. Sample · question 5 · Custom NER for domain entity types

    Quenby Aerospace needs to pull its own entity types, such as part numbers and maintenance codes, out of engineer reports. The prebuilt entity categories don't include these types. Which Azure Language feature should it use?

    • A.Prebuilt sentiment analysis
    • B.Custom named entity recognition trained on labeled examples from its reportscorrect
    • C.Language detection
    • D.Text-to-speech

    Why: Custom named entity recognition lets you train a model on your own labeled text to extract custom entity categories that prebuilt NER doesn't cover. Sentiment analysis measures opinions, language detection identifies the language, and text to speech produces audio.

    Open this question on its own page →
  6. Sample · question 6 · Text analytics for health on clinical notes

    A research hospital wants to extract medications, dosages, diagnoses and their relationships from unstructured clinical notes without building its own model. Which preconfigured capability of Azure Language handles this?

    • A.Conversational language understanding
    • B.Entity linking
    • C.Key phrase extraction
    • D.Text analytics for healthcorrect

    Why: Text analytics for health is a preconfigured feature that extracts and labels medical information, such as medications, dosages and diagnoses, and the relations between them, from unstructured clinical text. Conversational language understanding predicts user intents, entity linking connects mentions to a knowledge base, and key phrase extraction lists general topics.

    Open this question on its own page →
  7. Sample · question 7 · Field schema suggestion for new documents

    Ottery Archives has received a new type of shipping manifest it has never processed before. Before building a custom Content Understanding analyzer, the team wants the service to look at sample manifests and suggest which fields could be extracted. Which prebuilt analyzer is designed for this?

    • A.prebuilt-read
    • B.prebuilt-documentFieldSchemacorrect
    • C.prebuilt-audioSearch
    • D.prebuilt-receipt

    Why: prebuilt-documentFieldSchema is a utility analyzer that analyzes documents and proposes an appropriate field schema, which helps you discover the structure of a new document type before defining a custom analyzer. prebuilt-read performs basic OCR without suggesting fields, prebuilt-receipt extracts a fixed set of receipt fields, and prebuilt-audioSearch processes audio.

    Open this question on its own page →
  8. Sample · question 8 · Base analyzer for custom document analyzers

    A developer at Pellow Freight is creating a custom Content Understanding analyzer for scanned delivery notes. When defining the analyzer, which base analyzer should the baseAnalyzerId property reference?

    • A.prebuilt-audio
    • B.prebuilt-callCenter
    • C.prebuilt-documentcorrect
    • D.prebuilt-video

    Why: Custom analyzers inherit from one of four base analyzers, prebuilt-document, prebuilt-image, prebuilt-audio or prebuilt-video, chosen to match the content type, and they reference it through baseAnalyzerId. Delivery notes are documents, so prebuilt-document is the right parent. prebuilt-audio and prebuilt-video are for other modalities, and prebuilt-callCenter is a domain-specific analyzer rather than a base analyzer.

    Open this question on its own page →
  9. Sample · question 9 · Publishing agents to Teams and Copilot

    Stannard Partners has tested an HR agent in Foundry and now wants employees to use it where they already work, inside Microsoft Teams and Microsoft 365 Copilot. What should the team do?

    • A.Export the agent's instructions to a text file and email it to staff
    • B.Redeploy the underlying model as a Global Batch deployment
    • C.Publish the agent so it gets a stable endpoint and share it through Microsoft Teams and Microsoft 365 Copilotcorrect
    • D.Convert the agent into a Content Understanding analyzer

    Why: Foundry Agent Service supports publishing an agent as a managed resource with a stable endpoint and distributing it through channels such as Microsoft Teams and Microsoft 365 Copilot. Emailing instructions doesn't give users a working agent, Global Batch is for asynchronous bulk jobs, and an analyzer extracts content rather than hosting a conversational agent.

    Open this question on its own page →
  10. Sample · question 10 · Generating video from text prompts

    The marketing team at Velloway Cycles wants to create a short clip of "a cyclist riding through autumn woods at sunrise" from that text prompt and experiment with aspect ratio and duration in the Foundry portal. What should they use?

    • A.Speech translation
    • B.An embedding model in the model playground
    • C.A video generation model such as Sora 2 in the video playgroundcorrect
    • D.The prebuilt-videoSearch analyzer

    Why: Video generation models, such as Azure OpenAI Sora 2, create video from text prompts, and Foundry opens a video playground for them with controls such as aspect ratio and duration. An embedding model produces vectors, prebuilt-videoSearch analyzes existing videos rather than creating them, and speech translation works on audio.

    Open this question on its own page →
  11. Sample · question 11 · Agent-only guardrail intervention points

    Foundry guardrails scan for risks at defined intervention points. Which two intervention points apply only to agents and not to model deployments? (Select TWO.)

    • A.Tool callcorrect
    • B.Output
    • C.User input
    • D.Tool responsecorrect
    • E.Model training data

    Why: Guardrails support four intervention points: user input, tool call, tool response and output. Tool call, which covers what an agent proposes to send to a tool, and tool response, which covers what a tool returns to the agent, apply only to agents. User input and output apply to both models and agents, and training data is not a guardrail intervention point.

    Open this question on its own page →
  12. Sample · question 12 · Speaker diarization in transcripts

    Hartwell Clinic records two-person telehealth consultations and wants each transcript to show which participant said each sentence. Which Azure Speech capability provides this?

    • A.Speaker diarization during speech to text transcriptioncorrect
    • B.Pronunciation assessment
    • C.Neural text to speech with SSML
    • D.Key phrase extraction in Azure Language

    Why: Diarization separates the voices in an audio recording and labels each recognized phrase with a speaker, so the transcript shows who said what. Pronunciation assessment scores how accurately a speaker pronounces words, text to speech generates audio rather than transcribing it, and key phrase extraction finds the main topics in text without identifying speakers.

    Open this question on its own page →

Like the sample?

Other practice exams