Tuesday, 29 September 2026

Managing Database on ExaCC (X10M) part3



Part 1: Project Details & Infrastructure Architecture Blueprint
Project Context: Mission-Critical Tier-0 Core Banking Core Migration to ExaCC X10M
  • Objective: Migrate a highly concurrent, 80 TB legacy on-premises Oracle database environment into Exadata Cloud@Customer (ExaCC) X10M running Oracle Database 23ai/26ai containerized databases (CDB/PDB architecture) to enable AI vector searches for fraud detection, achieve sub-millisecond transactions, and streamline cloud operations using the OCI CLI.
  • Target Architecture:
    • ExaCC X10M Quarter Rack Elastic Configuration: High-performance database nodes utilizing AMD EPYC processors coupled with extreme performance storage servers running RoCE (RDMA over Converged Ethernet) network fabric.
    • Grid Infrastructure (GI) & RAC: 2-Node RAC cluster using multi-tenant configurations.
    • Storage Tier: ASM Flex Disk groups utilizing Exadata Smart Flash Cache, persistent memory (PMEM/XRMEM) accelerators, and Hybrid Columnar Compression (HCC). 

Part 2: Technical Interview Questions & Answers
Q1: How do the new micro-architectural differences in ExaCC X10M affect Grid Infrastructure and RAC cache fusion layer operations compared to older X9M systems?
Answer: ExaCC X10M replaces Intel processors with AMD EPYC processors, drastically scaling up the core count per database server. It expands the RoCE network fabric bandwidth to utilize extreme low-latency RDMA pathways directly to storage cells. 
  • RAC Cache Fusion Layer: With the massive core densities of X10M, inter-instance Global Cache Service (GCS) requests utilize hardware-assisted network virtualization and RoCE engine paths. This significantly reduces gc current block receive and gc cr block receive latencies below 150 microseconds under high concurrency workloads. 
  • GI Resource Tuning: Because of the core density, standard LMS process limits (CLUSTER_DATABASE_INSTANCES sizing) must scale dynamically. Oracle 23ai/26ai natively features self-tuning background processes to automatically allocate LMS processes based on the expanded CPU thread topology without overloading the OS scheduler.
Q2: In Oracle 23ai and 26ai, Lock-Free Reservations were introduced. How do you implement this feature in a highly concurrent RAC environment to eliminate application blocks?
Answer: Lock-Free Reservations solve the classic "hot-spot row lock" dilemma in high-throughput transactional environments (e.g., balance updates or inventory counters) where traditional locking mechanisms block concurrent updates across instances.
  • Mechanism: Instead of placing an exclusive row-level lock on the balance column, the database reserves the requested amount out of the total value logic using an internal reservation journal. This allows multiple sessions across RAC Node 1 and Node 2 to concurrently modify the same row without getting blocked by enq: TX - row lock contention.
  • Implementation: The column must be designated with the RESERVABLE keyword:
    sql
    ALTER TABLE bank_accounts MODIFY (account_balance RESERVABLE);
    

Q3: How do you perform automated, Zero-Downtime lifecycle operations on an ExaCC X10M multi-tenant database using the OCI CLI instead of the console interface?
Answer: Complex infrastructure operations on ExaCC should be managed through scripted pipelines using the OCI CLI. For instance, to scale up the OCPUs on an autonomous or co-managed ExaCC VM cluster seamlessly, or to provision a new Pluggable Database (PDB) within a 23ai Container Database (CDB): 
  • OCI CLI Provisioning Command:
    bash
    oci db pluggable-database create \
      --cdb-id ocid1.autonomousdatabase.oc1.eu-frankfurt-1.abx... \
      --pdb-name PDB_PROD_FRAUD \
      --admin-password "Complex_Pass_23ai" \
      --compartment-id ocid1.compartment.oc1..aaaaaaaax...
    
    Using the OCI CLI bypasses manual OCI Console interactions, allowing integration into Jenkins or OCI DevOps pipelines to dynamically scale resources during batch cycles without interrupting active RAC user connections. 

Part 3: Infrastructure Challenges & Real-World Failures
Challenge 1: ASM Split-Brain and Voting Disk Eviction under RoCE Network Partitioning
  • The Scenario: Due to a misconfigured switch top-of-rack (ToR) interface link aggregation on the ExaCC client/backplane network, transient packet drops occur on the RoCE network.
  • The Consequence: Nodes lose heartbeats over the interconnect. Grid Infrastructure enters a voting disk split-brain resolution scenario. 
  • The Resolution:
    1. Review cssdagent.log and ocssd.log to identify the missing node heartbeat.
    2. Query cluster synchronization statuses using crsctl check cluster -all.
    3. Validate RoCE health directly at the Exadata storage level using cell commands:
      bash
      cellcli -e LIST CELLATTRIBUTES name, interconnectOrder, status
      
      Isolate the broken link using ibdiagnet or native RoCE link tracing utilities to verify RDMA over Converged Ethernet link statuses.
Challenge 2: PDB Vector Search Memory Starvation in Oracle 23ai / 26ai
  • The Scenario: Developers heavily adopt AI Vector Search capabilities natively embedded within Oracle 23ai/26ai to process massive embeddings.
  • The Consequence: Vector data vectors are cached primarily in the SGA (VECTOR_MEMORY_AREA). A sudden surge of high-dimensional vector calculations exhausts the allocated database vector memory pools, throwing ORA-04031: unable to allocate bytes of shared memory errors and bottlenecking downstream transactional pipelines.
  • The Resolution:
    1. Dynamically expand the vector memory area within the database instance:
      sql
      ALTER SYSTEM SET VECTOR_MEMORY_AREA=16G SCOPE=BOTH SID='*';
      
      Implement an automated OCI Monitoring Alarm utilizing MQL expressions to track the database component metrics and trigger a serverless OCI Function via the CLI to adaptively scale up PDB memory profiles. 

Part 4: Step-by-Step Advanced Troubleshooting Playbook
Scenario: Intermittent Sub-Second Query Spikes on an ExaCC X10M 23ai Database
Follow this diagnostic hierarchy to debug application latency spikes on your engineered system:
[Step 1: Check Database Metrics via AWR]
                │
                ├──> Look for: 'gc current block busy' or 'cell smart table scan'
                │
[Step 2: Isolate Storage Component Layer via CellCLI]
                │
                ├──> Run: cellcli -e LIST ACTIVEREQUEST
                └──> Check for disk/flash disk latencies or IORM throttling
                │
[Step 3: Analyze Operating System & Grid Infrastructure Networks]
                │
                ├──> Run: traceroute -i bond0 <Storage_Cell_IP>
                └──> Check: oclumon dumpnodeview -allnodes
1. Analyze Database Bottlenecks (AWR & ASH)
Generate an ASH report during the exact timeframe of the spike. Look out for the following dominant wait events: 
  • cell smart table scan: Exadata Smart Scan offloading is functioning, but check if high volume is overwhelming disk modules. 
  • gc current block busy / gc cr block busy: Indicates contention in the RAC Global Cache layer. The block is being modified on another node while a remote node requires it. 
2. Interrogate the Exadata Storage Servers (CellCLI)
Log into the database server nodes and execute remote calls across the cells: 
bash
# Check for long-running I/O operations directly at the cell layers
cellcli -e "LIST ACTIVEREQUEST WHERE ioReason = 'Smart Scan'"

# Check if IORM (I/O Resource Management) is actively throttling the PDB database workloads
cellcli -e "LIST METRICCURRENT WHERE name LIKE 'CL_BY_GE_DIR_W'"
3. Diagnose the RoCE Network Interconnect
If the wait events point to network lag, analyze the interconnect statistics at the OS layer via GI cluvfy and oclumon:
bash
# Dump real-time cluster node view information
oclumon dumpnodeview -allnodes -v

# Validate the integrity of the cluster network components
cluvfy comp nodecon -n all -verbose
Part 5: Comprehensive Verification & Test Case
Test Case Objective
Verify that the Oracle 23ai Lock-Free Reservations feature works optimally on an ExaCC X10M 2-Node RAC Cluster under multi-instance concurrent transactions without triggering resource block locks or transactional deadlocks.
       [NODE 1]                                      [NODE 2]
   Session A (T1)                                Session B (T2)
         │                                             │
         ▼                                             ▼
  Reserve $100 from                             Reserve $50 from
  Account Balance                               Account Balance
         │                                             │
         └───────────────► [SHARED DATA] ◄─────────────┘
                     account_balance RESERVABLE
                                 │
                                 ▼
                     SUCCESS: Both Commit Without 
                        Row Lock Contention!
Step-by-Step Test Execution Plan
1. Setup Phase (Executed via Node 1)
Create a test database table inside the pluggable database (PDB_PROD_FRAUD) containing a designated RESERVABLE numeric column.
sql
ALTER SESSION SET CONTAINER = PDB_PROD_FRAUD;

CREATE TABLE customer_ledger (
    customer_id   NUMBER PRIMARY KEY,
    customer_name VARCHAR2(100),
    balance       NUMBER RESERVABLE CONSTRAINT min_balance CHECK (balance >= 0)
);

INSERT INTO customer_ledger VALUES (999, 'Enterprise Account Corp', 50000);
COMMIT;
2. Concurrent Transaction Phase (Simulating Multi-Node RAC Loads)
  • Node 1 - Session A (Time T1): Deduct an amount from the account balance without issuing an intermediate commit.
    sql
    -- Executing on RAC Instance 1
    UPDATE customer_ledger 
    SET balance = balance - 10000 
    WHERE customer_id = 999;
    -- Note: Do NOT issue a commit yet!
    

  • Node 2 - Session B (Time T2): Concurrently deduct another amount from the same customer record using the second RAC instance.
    sql
    -- Executing on RAC Instance 2
    UPDATE customer_ledger 
    SET balance = balance - 5000 
    WHERE customer_id = 999;
    

3. Expected Evaluation Matrix
Metric / Observed OutcomeStandard Database (Legacy Mode)Oracle 23ai/26ai Reservable Mode
Session B Execution StateBlocked (Hangs waiting for TX Row Lock)Immediate Success (No block encountered)
Dominant Cluster Wait Eventenq: TX - row lock contentionNone / Normal CPU Execution
Final State (After both COMMIT)Balance updates sequentiallyBalance updates concurrently instantly
4. Validation Script
Execute a tracking query across both sessions to verify that updates were journaled properly prior to committing changes:
sql
SELECT customer_id, balance, USERENV('Instance') AS instance_id FROM customer_ledger WHERE customer_id = 999;
-- Issue commit on both nodes to flush journal changes down to the base table segments
COMMIT;



 Part 1: Project Overview (The Context)

Project Detail: Enterprise Infrastructure Hardening & Compliance Migration
  • Infrastructure: Exadata Cloud@Customer (ExaCC) X10M (Quarter/Half Rack configurations).
  • Objective: Remediate critical CVEs (Operating System, Grid Infrastructure, Database, RoCE Network switches) and meet strict federal/corporate regulatory compliance targets without disrupting high-throughput, mission-critical online transaction processing (OLTP) and data warehousing workloads.
  • The Split-Responsibility Challenge: On ExaCC X10M, Oracle manages the physical hardware, Dom0 (Hypervisor), RoCE network switches, and Power Distribution Units (PDUs). The customer is entirely responsible for the DomU (User Virtual Machines), including guest OS patching, database software security, IAM compliance, encryption keys, and localized auditing configurations.

Part 2: Interview Questions and Answers (Q&A)
Q1: How do you handle vulnerability remediation on ExaCC X10M given Oracle's shared responsibility model?
Answer: Remediation is executed in two parallel streams. Oracle automatically handles infrastructure updates (Hypervisor, Storage Cells, RoCE switches) based on scheduled maintenance windows via the Oracle Cloud Infrastructure (OCI) Control Plane. As the ExaCC Administrator, my responsibility covers the DomU layer. I leverage ExaCLI, Patch Manager (patchmgr), and the OCI Console/API to orchestrate rolling patch applications across the Grid Infrastructure (GI) and Database homes, ensuring zero downtime by leveraging Oracle RAC rolling upgrades.
Q2: How does the ExaCC X10M architecture change your approach to isolating and auditing network traffic for compliance?
Answer: X10M relies on Secure RDMA Fabric Isolation (Secure Fabric) over RoCE. Unlike traditional environments where standard network taps are used, X10M separates traffic at the hardware level between tenant VMs. For compliance auditing, we configure internal firewall rules (iptables / firewalld) inside the DomU, implement Oracle Connection Manager (CMAN) to log and audit incoming traffic paths, and use Oracle Unified Auditing to capture database-level interactions.
Q3: An external compliance audit flags that "root" or broad "sudo" access on ExaCC components violates the principle of least privilege. How do you remediate this?
Answer: On ExaCC X10M, full root access to the underlying infrastructure is restricted. To bridge this audit gap, we enforce Oracle Database Vault to separate the duties of the Cloud Administrator from the Data Owner. Within the DomU OS, we restrict sudo privileges using fine-grained rule definitions in /etc/sudoers.d/ and integrate the environment with a Privileged Access Management (PAM) tool like CyberArk or HashiCorp Vault to generate ephemeral, fully audited access keys.
Q4: How do you validate that a newly applied quarterly Security Patch Update (SPU/RU) hasn't broken Exadata-specific optimizations like Smart Scan or Hybrid Columnar Compression (HCC)?
Answer: I run exacheck immediately post-patching to verify that the software versions across DomU and the storage cells are fully compatible. To validate functional performance optimization, I execute a baseline diagnostic query using AUTOTRACE or analyze V$SYSSTAT metrics to ensure that system statistics like cell physical IO bytes eligible for smart scan are incrementing correctly.

Part 3: Operational Challenges & Troubleshooting
1. Mismatched Infrastructure vs. Guest VM Versions
  • The Challenge: Oracle auto-updates the Exadata Storage Servers to a newer release (e.g., Exadata System Software 26.1+), but the internal guest VM (DomU) is running an older Release Update (RU). This skew can trigger unexpected ORA errors or bypass critical security fixes.
  • Troubleshooting:
    • Execute imageinfo on the storage layer via OCI CLI and compare it with cat /opt/oracle.cellos/iso_version inside the DomU.
    • Deploy exacheck with the explicit flag to check patch dependencies. If a conflict is present, immediately schedule a rolling DomU patch window using patchmgr.
2. OCI Control Plane Connectivity Loss During Security Hardening
  • The Challenge: Strict corporate firewalls or aggressive local iptables rule updates inside the DomU can inadvertently block the OCI management agent, breaking the hybrid cloud control plane connection.
  • Troubleshooting:
    • Ensure that ports 443 (HTTPS) and 1522 (default scan listener tracking) are open to the specific OCI infrastructure CIDR blocks.
    • Check the status of the OCI management agent: systemctl status mgmt_agent.
    • Review log outputs at /var/lib/oracle-cloud-agent/ to discover dropped packets or certificate blockages.
3. Audit Logging Storage Bottlenecks
  • The Challenge: Enabling comprehensive Oracle Unified Auditing and OS-level syslogs to fulfill compliance mandates generates massive I/O overhead, threatening to saturate local VM storage disks (/u01).
  • Troubleshooting:
    • Implement Oracle Audit Vault and Database Firewall (AVDF) to continuously offload audit trails from the local ExaCC environment into a centralized repository.
    • Configure dynamic log rotation schedules (/etc/logrotate.conf) and map massive audit tables to a dedicated tablespace hosted on an Exascale shared volume or high-performance ASM disk group.

Part 4: Test Case Template
This practical test case demonstrates how to apply, verify, and pass an audit check for a Critical Security Patch installation on an ExaCC X10M cluster.
Test Case IDTC-EXACC-SEC-004
Test TitleQuarterly Release Update (RU) Vulnerability Remediation & Audit Validation
ComponentDomU (Guest OS), Oracle Grid Infrastructure, Oracle Database
Prerequisites1. Access to OCI Console with Exadata Infrastructure Admin rights.
2. Latest Patch Bundle downloaded to the staging area.
3. Baseline performance metrics recorded via AWR.
Execution Steps1. Execute pre-patch compliance checks: Run ./exacheck and log all initial findings.
2. Apply the Grid Infrastructure and Database patches in a rolling fashion using the OCI Console or ./patchmgr -dbnode <node_list> -action patch.
3. Once complete, query the active software inventory: opatch lsinventory.
4. Run the post-patch compliance suite: ./exacheck --profile security.
Expected Result1. All target cluster nodes report successful patch installation status.
2. Databases and listeners remain online throughout the rolling update cycle.
3. opatch lsinventory correctly lists the newly targeted CVE identifiers.
4. The post-patched exacheck report shows 0 "Critical" failures.
Audit Evidence1. Text log outputs generated by opatch lsinventory.
2. Timestamped PDF/HTML performance compliance reports from exacheck.

Project Details: Large-Scale Financial Cloud Migration
  • Platform & Infrastructure: Oracle Exadata Cloud at Customer (ExaCC) X10M Multi-Node RAC Cluster running Oracle Database 19c Enterprise Edition (Extreme Performance).
  • Architecture Strategy: Multi-tenant Container Databases (CDB/PDB) utilizing United Mode for unified key infrastructure across shared clusters.
  • Security Standard: Strict AES-256 Tablespace Encryption mandated for all at-rest tablespaces (CLOUD_ONLY).
  • Key Storage Scheme: Transitioned from basic file-based Auto-Login software wallets (ewallet.p12 and cwallet.sso) stored in the shared WALLET_ROOT/tde/ directories to centralized enterprise management via Oracle Key Vault (OKV).

Interview Questions & Answers
Q1: How do you handle TDE keystores across multiple nodes on an ExaCC X10M RAC environment?
Answer: On ExaCC X10M RAC clusters, TDE software wallets must be stored on a shared filesystem—typically Oracle ASM or ACFS under the WALLET_ROOT/tde/ path—allowing all active cluster nodes simultaneous access. Individual local wallets per node are completely unsupported. For enterprise scale, we configure sqlnet.ora and TDE_CONFIGURATION parameters to bind database endpoints directly to an external network HSM or Oracle Key Vault (OKV), maintaining high-availability synchronization via automated endpoints.
Q2: What is the risk of using OCI tooling (dbaascli) versus standard SQL commands when mutating TDE wallets?
Answer: If you use low-level SQL commands (ADMINISTER KEY MANAGEMENT...) to alter wallet configurations or switch states out-of-band, you risk fracturing the Cloud Control Plane state. When the ExaCC automated tooling tries to perform maintenance operations (like database patching, point-in-time recovery, or scale-out), it reads the OCI registry metadata. If the physical wallet path or password mismatches the cloud registry, cloud updates will fail hard, potentially locking down automated backup actions. The safest approach is always executing modifications via dbaascli database ... or the OCI Console interface wherever natively supported.
Q3: Explain the difference between an auto-login wallet and a standard software wallet, and why both matter in ExaCC.
Answer: The standard software wallet (ewallet.p12) is password-protected and is required whenever performing administrative write actions, such as rotating a master encryption key or adding new credentials. The auto-login wallet (cwallet.sso) is derived directly from the standard wallet and allows the database instances to automatically read the Master Encryption Key (MEK) at server boot time without requiring physical human entry of a password. This is essential on ExaCC environments to ensure high availability during unattended cluster node reboots or automated patch applications.

Operational Challenges & Troubleshooting
1. Cloud Infrastructure vs. Database State Mismatch
  • Challenge: The team runs a master key rotation via explicit SQL commands inside the PDB, but subsequent automated backups triggered through the OCI / Cloud@Customer console crash with access validation errors.
  • Root Cause: Cloud tooling keeps its own tracking record of the wallet password and status. When performing keys mutations via native SQL, the orchestration layer loses alignment.
  • Resolution: Sync changes or run maintenance workflows directly using the dbaascli utility wrapper on the ExaCC compute nodes to guarantee that internal operational metadata syncs uniformly with the OCI control registry.
2. Race Conditions on Auto-Login Wallets in Dynamic Clusters
  • Challenge: During high-velocity database initialization or parallel patching cycles across cluster nodes, intermittent ORA-28374: typed master key not found in wallet errors trigger.
  • Root Cause: The cwallet.sso file gets updated unevenly across multi-node shared disk architectures if permissions or file-locking states delay local cluster cache sweeps.
  • Resolution: Verify permissions on the shared cluster path (chmod 600 for the oracle user). Avoid hard-copy steps across individual instances. Instead, ensure the operational runtime leverages standard parameter declarations:
    sql
    ALTER SYSTEM SET TDE_CONFIGURATION="KEYSTORE_CONFIGURATION=FILE" SCOPE=BOTH;
    


Test Case: End-to-End Master Key Rotation Validation
Test IDObjectiveSteps to ExecuteExpected ResultsPass/Fail Criteria
TC-TDE-01Validate zero-downtime TDE Master Encryption Key (MEK) Rotation1. Query current key state from v$encryption_keys.
2. Log into database cluster control node via dbaascli.
3. Trigger the key rotation command:
dbaascli database rotateKey --dbName EXAPROD
4. Re-verify the active database key view.
1. An entirely new master key identifier is generated.
2. Active application user queries continue to run concurrently without connectivity loss.
Pass: Data continues to stream out without error; new key entries reflect correctly inside v$encryption_keys.

Project Scenario & Context
During an enterprise migration and consolidation project, Oracle Exadata Cloud at Customer (ExaCC) X10M was deployed to host critical workloads. Before executing a major change—such as a Quarterly Infrastructure Patching (Grid Infrastructure/OS/Firmware upgrade) or severe architectural modifications—a rigorous maintenance protocol mandates executing Autonomous Health Framework (AHF) / EXAchk health checks immediately before and after the change.
This guarantees baseline configuration compliance, uncovers pre-existing underlying risks, and verifies that the system has safely returned to a high-availability state without introducing configuration drifts.

Interview Questions & Answers
Q1: Why is running EXAchk mandatory both before and after a maintenance window on ExaCC X10M?
A: Running it before establishes an environmental baseline and catches existing faults (e.g., cell disk alerts, skewed configurations, or grid infrastructure bugs) that could crash the update process. Running it after ensures no configuration drifts occurred, verifies that software versions match perfectly across the nodes, and uses the EXAchk Diff utility (-diff) to isolate exact changes introduced during maintenance.
Q2: How does executing EXAchk differ on ExaCC X10M compared to an on-premises Exadata machine?
A: On ExaCC, the infrastructure layer (Dom0, Storage Cells, and RoCE Network Switches) is managed exclusively by Oracle. Customers have root access to the DomU (virtual machines). When you run exachk from a DomU database node, it communicates with the underlying storage layer using standard exacli calls rather than direct SSH passwordless access to root cells, adapting to ExaCC's strict security boundaries.
Q3: Which explicit flags are used to run upgrade readiness checks using AHF/EXAchk?
A: To run pre-upgrade compliance checks, use ./exachk -u -o pre. 
To validate system state post-upgrade, execute ./exachk -u -o post.
Q4: If an EXAchk process hangs indefinitely on an X10M compute node during a pre-check, how do you diagnose and circumvent it?
A: The AHF/EXAchk architecture has a built-in watchdog process that terminates hung actions based on internal timers. If it hangs, you look into output_dir/log/exachk.log to find the exact component causing the stall. To circumvent network or switch response latency, you can bump the environment variables like RAT_TIMEOUT or execute exachk -local to run it solely on the local compute node, subsequently using -merge to combine reports from other nodes.

Key Operational Challenges
  • Asymmetric Security Configurations: ExaCC X10M utilizes RoCE (RDMA over Converged Ethernet) network fabrics instead of traditional InfiniBand. Because storage cells block direct SSH logins from the customer side, security parameters sometimes prevent the standard EXAchk engine from polling storage health metrics.
  • Timeouts on Large Scale Consolidations: When multiple databases are consolidated onto an X10M rack, default execution times for automated collection can time out while polling deep software stack metrics across clustered instances.
  • Outdated AHF Engine Engines: Exadata rules change rapidly. Running an old version of exachk creates high false-positive rates by checking deprecated parameters against the state-of-the-art X10M hardware architecture.

Troubleshooting Playbook
Issue / ErrorRoot CauseTarget Remediation
Hangs at Storage/Cell VerificationRestrictive exacli connectivity rules or high network latency.Set export RAT_PASSWORDCHECK_TIMEOUT=40 or execute temporary storage cell unlocks using appropriate AHF flags.
Insufficient Space in System DirectoriesThe local directory or /tmp has run out of space during file compression.Redirect the workspace folder by setting the environment variable export RAT_TMPDIR=/u02/app/oracle/tmp before starting the script.
Discovery Failures (Missing DB/ASM targets)Environment profiles fail to accurately resolve specific Grid Infrastructure paths.Explicitly enforce location indicators prior to execution, such as export RAT_CRS_HOME=$GRID_HOME and export RAT_ASM_HOME=$ASM_HOME.

Detailed Test Case: Executing Before & After Changes
Objective
Successfully baseline an ExaCC X10M system environment, execute a minor rolling Grid Infrastructure change, evaluate health status post-change, and isolate differences.
Step 1: Pre-Change Baseline Execution
Execute a thorough check across all cluster instances from the primary database cluster node:
bash
# Log in as root or GI Owner on Node 1
cd /opt/oracle.ahf/exachk

# Verify the AHF framework tool status
ahfctl statusahf

# Run the complete health check engine pre-change
./exachk -a

  • Expected Result: An HTML report is successfully outputted to the local directory. Review the report for any CRITICAL/FAIL flags that must be resolved prior to the scheduled maintenance window.
Step 2: Implement System Changes
(Execute the planned patching cycle, parameter change, or infrastructure alteration.)
Step 3: Post-Change Verification Execution
Rerun the validation program cleanly to gather the post-maintenance state:
bash
./exachk -a

  • Expected Result: A new matching comprehensive verification report is generated.
Step 4: Differential Mapping & Isolation
Compare both execution files using the direct differential utility framework:
bash
./exachk -diff <path_to_pre_change_zip> <path_to_post_change_zip>

  • Expected Result: A tailored Diff Report is created. Review this output carefully to confirm that no unauthorized configuration parameters were modified and that software builds match accurately across the entire environment.

Project Detail (The Context)
Project Name: Mission-Critical Core Banking & Analytics Migration to ExaCC X10M
Environment: Exadata Cloud@Customer X10M Quarter Rack (scalable up to Multi-Rack).
Database Profile: Multi-terabyte Oracle 19c/23ai Container Databases (CDB/PDB) running hybrid workloads (OLTP & Data Warehouse/Analytics).
Key Features Utilized:
  • AMD EPYC processors (high core density per database server).
  • Exadata RDMA Memory (XRMEM) replacing traditional Flash Cache for ultra-low latency reads.
  • PCIe Gen 5 NVMe Flash for high-throughput Smart Scans.
  • RoCE (RDMA over Converged Ethernet) 100 Gbps internal network fabrics.