Google Cloud Professional Data Engineer · Free practice question 4 of 12
Data Validation Tool after migration
After moving 400 tables from an on-premises PostgreSQL warehouse to BigQuery, Holloway Insurance's auditors want evidence that every table arrived complete and unchanged, including row counts, column sums and row-level comparisons. Which approach provides this with the least custom code?
- A.Compare the storage size of each table in INFORMATION_SCHEMA.TABLE_STORAGE with the source
- B.Accept the migration job's success status as proof of completeness
- C.Run Google's open-source Data Validation Tool to compare counts, aggregates and row hashes between source and target
- D.Run Knowledge Catalog (formerly Dataplex Universal Catalog) data profile scans on the BigQuery tables
Show answer and explanation
Correct answer: C. Run Google's open-source Data Validation Tool to compare counts, aggregates and row hashes between source and target
Why: The Data Validation Tool connects to both systems and runs column, row-count, aggregate and row-hash validations, producing a report of matches and differences. Storage size differs between engines because of compression and formats. A job's success status does not prove the data matches, and profiling the target alone has nothing to compare against.
More free Google Cloud Professional Data Engineer questions
- Sliding windows for moving averages
- Turbo replication on dual-region buckets
- Assured Workloads for sovereign controls
- Pub/Sub Cloud Storage subscription archive
- Bigtable garbage collection by age
- Search indexes for log lookups
- Data Studio viewer's credentials
- Integer-range partitioning on an ID
- Approximate distinct counts for dashboards
- Scheduled queries for a single SQL job
- Pub/Sub message storage policy regions