Tuesday, 6 October 2026

Exadatacc.x10m.xl Architecture Detail and Troubleshooting Steps

Oracle Exadata Database Machine X10M


Question : There is 4 node database exadatacc.x10m.xl . there is intermittent connection on mainly on node4from app server


Phase 1: Verify RAC VIP & Listener Health on Node 4
When applications connect via SCAN, the SCAN listener redirects the client to the local VIP listener of a specific node. If Node 4's listener or VIP drops packets, clients encounter intermittent timeout or handshake errors (such as ORA-12170 or ORA-12541). 
  1. Check the Local Listener and VIP Status:
    bash
    crsctl status resource ora.LISTENER.lsnr -p
    crsctl status resource ora.LISTENER.lsnr -v
    srvctl status vip -node <node4_hostname>
    

  2. Review Listener and Connection Rates on Node 4:
    • Inspect the listener log:
      bash
      $ORACLE_BASE/diag/tnslsnr/<node4_hostname>/listener/trace/listener.log
      
      Look for TNS-12535 (timeout), connection refused errors, or registration drops.
    • If connection rates are very high, ensure the listener connection backlog isn't exhausted by verifying RATE_LIMIT or increasing QUEUESIZE in listener.ora (e.g., QUEUESIZE=1024).

Phase 2: Check Physical Client Network & Bonding (bondeth0)
Exadata X10M compute nodes utilize aggregated client interfaces (typically bondeth0 using LACP / 802.3ad). A flapping physical link or bad SFP on Node 4 will cause intermittent packet loss. 
  1. Check Bonding Interface Status:
    bash
    cat /proc/net/bonding/bondeth0
    
    • Confirm both slave interfaces (e.g., eth0, eth1 or RoCE/Ethernet client ports) show MII Status: up.
    • Ensure LACP state is synchronized and not constantly re-negotiating. 
  2. Inspect for Interface Errors and Drops:
    bash
    ip -s link show bondeth0
    ethtool -S <slave_interface_name> | grep -E "err|drop|fifo|frame"
    

  3. Trace Route & Latency from the App Server:
    • Run continuous pings with MTU matching your network from the application server specifically to Node 4's VIP and Physical IP:
      bash
      ping -c 100 -s 1472 <node4_vip_ip>
      
      Check for dropped packets or variable RTT.

Phase 3: Check Instance Process & Session Limits
If the local instance on Node 4 hits process limits, incoming connections will hang or get rejected intermittently.
  1. Check Resource Utilization:
    sql
    SELECT inst_id, resource_name, current_utilization, max_utilization, initial_allocation 
    FROM gv$resource_limit 
    WHERE resource_name IN ('processes', 'sessions') AND inst_id = 4;
    

  2. Review Alert Log for Hangs or Evictions:
    • Inspect the database alert log on Node 4:
      bash
      $ORACLE_BASE/diag/rdbms/<db_name>/<instance4_name>/trace/alert_<instance4_name>.log
      
      Look for memory thrashing, HugePages exhaustion, CPU starvation warnings, or dead connection detection (DCD) kills.

Phase 4: Validate DNS and VIP Resolution
In multi-node RAC environments, inconsistent reverse DNS lookup or duplicate IP allocation can cause intermittent connections to fail on a single node: 

  • From the application server, verify that the Node 4 hostname, VIP, and SCAN addresses resolve consistently across all configured DNS servers:
    bash
    nslookup <node4_vip>
    nslookup <node4_hostname>
    

  • Verify that no IP conflicts exist for Node 4's VIP using arping -I <interface> <node4_vip>. 




or


tart

Based on the provided error logs and architecture context, restarting the Virtual IPs (VIPs) and SCAN listeners is a solid initial troubleshooting step, but it might not resolve the underlying root cause.
The logs indicate a mismatch between network layer drops and listener service registration status specifically impacting Node 4. Below is an architectural breakdown of why this is happening and a targeted action plan.
Error Analysis & Intermittent Root Causes
  1. TNS-12514: listener does not currently know of service requested... 
    • What it means: The application is attempting to connect to a specific service name, but the listener on the target node does not have that service registered in its dynamic service list.
    • Why it happens intermittently on Node 4: In an Oracle RAC/Exadata environment, instances dynamically register their services with the local and SCAN listeners via the PMON or LREG background processes. If Node 4 is experiencing high CPU spikes, internal resource contention, or inter-instance cluster interconnect delays, database service registration updates to the listeners will drop out or time out. 
  2. TNS-12535 / TNS-00505: Operation timed out 
    • What it means: The transport layers successfully initiated a connection handshake, but it timed out before completion.
    • Why it happens here: This typically points directly to a network firewall dropping idle database connection pools, asynchronous routing issues across Node 4's specific network interface card (NIC), or database listener queue saturation (backlog queue full). [

Step-by-Step Action Plan
Step 1: Execute the Planned Cluster Actions
Go ahead with your plan to bounce the VIP and the Scan Listener components to refresh the cluster resource routing layers. Run these via Grid Infrastructure (GRID_HOME):
bash
# Relocate/Restart Node 4 VIP (Forces network routing to refresh)
srvctl stop vip -n <node4_name> -f
srvctl start vip -n <node4_name>

# Check and update service registration status across listeners
srvctl status scan_listener
# If needed, restart the specific scan listener instance handling Node 4 routes
srvctl stop scan_listener -i <listener_number>
srvctl start scan_listener -i <listener_number>
Step 2: Force Instant Service Registration
Once the listeners are back up, log into the Oracle instance on Node 4 and manually force the database instance to re-register its services with the cluster listeners immediately:
sql
ALTER SYSTEM REGISTER;
Step 3: Check Firewall & Dead Connection Detection (DCD)
Your sqlnet.ora profile configuration (shown in Image 4) shows SQLNET.EXPIRE_TIME=10 is active. This is ideal because it instructs the database to send a keepalive probe every 10 minutes to prevent firewalls from terminating idle application connection pools. 
  • Verification: Check if there is an aggressive network security rule or a firewall sitting between the Application Server subnet and the Node 4 VIP/SCAN subnet that drops connections inside a tighter window (e.g., less than 10 minutes).
Step 4: Verify Dynamic Cross-Registration Settings
Ensure Node 4 can talk back cleanly to all SCAN listeners. Validate the initialization parameters on Node 4's instance:
sql
SHOW PARAMETER local_listener;
SHOW PARAMETER remote_listener;
  • Fix Requirement: remote_listener must point to your network SCAN alias so it can dynamically communicate availability metrics across the entire 4-node cluster.


or

To safely execute your action plan for restarting the node VIPs and SCAN listeners on an Oracle Exadata Cloud@Customer (ExadataCC) 4-node cluster, you must perform the operations via the Oracle Grid Infrastructure (grid) user using the Server Control (srvctl) utility. 
 an application error logging a delayed_connect_error_Connection_refused on port 1521 (pointing directly to a listener endpoint issue), target node 4 specifically while ensuring cluster services remain healthy.

Step-by-Step Action Plan
1. Pre-Check (Verify Current Status)
Before restarting any resources, check the exact layout and running status of your Node VIPs and SCAN listeners across all 4 nodes: 
bash
# Log in to node4 as the 'grid' user

# Check the status of all Node VIPs
srvctl status vip -proto

# Check the configuration and location of the SCAN listeners
srvctl status scan_listener
srvctl config scan
Note: In a standard setup, there are 3 SCAN VIPs/listeners distributed across the 4 nodes. Verify if one of the SCAN listeners is currently hosted on Node 4.
2. Restart Node 4 VIP
Restarting the node's Virtual IP will briefly drop and re-initialize the interface on that host.
bash
# Stop the VIP resource on Node 4
srvctl stop vip -n <node4_hostname> -f

# Verify it is offline or has stopped
srvctl status vip -n <node4_hostname>

# Start the VIP resource back up on Node 4
srvctl start vip -n <node4_hostname>
3. Restart the SCAN Listener
Because SCAN listeners depend entirely on SCAN VIPs, restarting a specific SCAN listener should be done using its instance/ordinal number (-i), or you can cycle the global SCAN listener service if a brief failover window is permitted: 
Option A: Restart only the specific SCAN listener instance running on Node 4
If srvctl status scan_listener showed that instance 2 or 3 is running on node 4, target just that ID: 
bash
# Stop the specific SCAN listener instance (e.g., instance 2)
srvctl stop scan_listener -i <instance_number>

# Start the SCAN listener instance back up
srvctl start scan_listener -i <instance_number>
Option B: Rolling Restart of all SCAN Listeners
If you prefer to clear and cycle the entire SCAN pool across the cluster:
bash
srvctl stop scan_listener
srvctl start scan_listener
4. Post-Check Verification
Verify that all services have successfully bound back to port 1521 and are accepting cluster registrations: 
bash
# Check clusterware resource summary status
crsctl status resource -t

# Confirm SCAN listeners are running normally
srvctl status scan_listener
 Important Cloud@Customer Guardrails
  • Do Not Use OS Reboot Commands: Avoid running Linux-level network interface restarts (ifdown/ifup) or service restarts directly via systemd (systemctl restart), as Oracle Grid Infrastructure will interpret this as an unexpected failure and trigger a node eviction or automated failover. Always use srvctl. 
  • Database Connection Impact: Stopping a Node VIP or SCAN listener will terminate active client connections that are currently routing through that specific component. Since connections are intermittent, ensure this is performed during a designated maintenance window or low-traffic period. 


or


Restarting the Node VIP and SCAN listeners is a valid tactical mitigation, but it carries immediate connection-failover impacts and will not resolve underlying network, listener process, or resource exhaustion issues.
Below is an assessment of your proposed plan, a structured step-by-step execution procedure, and root-cause validation steps to perform beforehand.
Pre-Action Health Check & Triage
Before restarting resources, capture state information to pinpoint why Node 4 specifically experiences intermittent drops:
  • Verify VIP & Listener Status: Check which node currently hosts the SCAN listeners and verify the state of Node 4’s local listener:
    bash
    srvctl status scan_listener
    srvctl status listener -node <node4_hostname>
    srvctl status vip -node <node4_hostname>
    

  • Inspect Listener Logs: Check $ORACLE_BASE/diag/tnslsnr/<node4_hostname>/<listener_name>/trace/ for TNS-12535 (timeout), TNS-12518 (listener cannot hand off), or OS errors.
  • Check Node 4 Resource Saturation: Ensure the OS has not hit socket, process, or file descriptor limits:
    bash
    # Check system process count and TCP socket states
    ps -ef | wc -l
    ss -s
    netstat -s | grep -i "listen"
    

  • Verify Network & MTU Consistency: Test latency and packet drops from the application server to both the SCAN VIPs and Node 4’s client VIP:
    bash
    ping -s 1472 -M do <node4_vip_ip>
    traceroute -T -p 1521 <node4_vip_ip>
    
    Impact Analysis
  • Local VIP Restart on Node 4: Any new incoming connection directed to Node 4 VIP will be temporarily rejected or failed over by Grid Infrastructure. In RAC, dropping a VIP will trigger a failover address notification, but in-flight uncommitted transactions on Node 4 that experience TCP resets may terminate depending on Application Continuity (TAC/FAN) configuration.
  • SCAN Listener Restart: SCAN listeners are distributed across the cluster (usually 3 SCAN IPs/listeners across nodes). Restarting SCAN listeners forces SCAN VIPs to migrate or reload, causing momentary connection retry spikes for all application servers connecting cluster-wide.
Recommended Step-by-Step Action Plan
Execute this during an approved maintenance or low-traffic window:
Phase 1: Graceful Local Listener & VIP Restart on Node 4
  1. Relocate or Stop Node 4 Local Listener:
    bash
    # As grid user:
    srvctl stop listener -listener LISTENER -node <node4_hostname>
    

  2. Stop and Start Node 4 VIP:
    bash
    srvctl stop vip -vip <node4_vip_name> -force
    srvctl start vip -vip <node4_vip_name>
    

  3. Start Node 4 Local Listener:
    bash
    srvctl start listener -listener LISTENER -node <node4_hostname>
    

  4. Register Database Services:
    sql
    -- On Node 4 database instance:
    ALTER SYSTEM REGISTER;
    
    Phase 2: SCAN Listener Restart (Cluster-Wide)
If the SCAN listener that routes to Node 4 needs recycling:
  1. Identify Node Hosting the Impacted SCAN:
    bash
    srvctl status scan_listener
    

  2. Restart SCAN Listeners Sequentially:
    bash
    # Cycle SCAN listeners one at a time to prevent cluster-wide blackout:
    srvctl stop scan_listener -i 1
    srvctl start scan_listener -i 1
    
    srvctl stop scan_listener -i 2
    srvctl start scan_listener -i 2
    
    srvctl stop scan_listener -i 3
    srvctl start scan_listener -i 3
    
    Post-Restart Verification
  • Check CRS Services: Run crsctl stat res -t and confirm all ora.LISTENER.lsnr, ora.*.vip, and ora.LISTENER_SCAN*.lsnr resources report ONLINE/ONLINE.
  • Validate App Connectivity: Run a looped SQL*Net connection test directly from the application server targeting Node 4 VIP and the SCAN address:
    bash
    for i in {1..20}; do nc -zv -w 2 <node4_vip_ip> 1521; sleep 1; done

or

 the error logs show an intermittent delayed_connect_error Connection_refused when attempting to connect to port 1521 on the database node IP.

Because the issue occurs intermittently and is heavily isolated to Node 4 of your 4-node Exadata Cloud@Customer (ExadataCC X10M XL) environment, it indicates that the core database network routing is likely fine, but Node 4 is specifically failing to handle or receive the connection requests.
You can systematically isolate and troubleshoot this issue on Node 4 through the following areas:
1. Check Local Listener & SCAN Status
If the listener or local Virtual IP (VIP) on Node 4 is down, unresponsive, or hanging, connections routed to it will return a connection refused error.
  • Log into Node 4 as the grid user and check the status of the local listener:
    bash
    srvctl status listener
    lsnrctl status
    
    Check if the VIP for Node 4 is properly running and bound to the interface:
  • bash
    srvctl status vip -node <node4_name>
    
    Verify the overall cluster and SCAN status across all nodes using Oracle Grid Infrastructure tools: 
  • bash
    crsctl check cluster -all
    srvctl status scan_listener
    
    2. Verify Port & Process Availability
An intermittent refusal often points to listener resource exhaustion or local operating system constraints closing the port.
  • On Node 4, check if the listener process is hitting system process limits or thread ceilings, which will cause it to drop incoming TCP handshakes.
  • Review the listener log (listener.log) specifically on Node 4 at the exact timestamps matching your application errors to see if there are corresponding TNS-12518, TNS-12514, or TNS-12502 errors.
  • Run a continuous tnsping or nc/telnet loop from the application server targeting Node 4's specific VIP address to check if the refusal drops synchronously with local resource spikes.
3. Review OS Firewall & Security Rules (IPTables / Firewalld)
An intermittent failure can happen if there is an IP conflict, or if an automated security daemon (like fail2ban or aggressive profile filtering) on Node 4 is dynamically blocking the application server's IP address.

  • Check the local firewall status on Node 4:
    bash
    sudo firewall-cmd --state
    # or review active rules:
    sudo iptables -L -n -v
    
    4. Check for Subnet/IP Asymmetry or RoCE Interconnect Issues
Since this is an Exadata X10M platform, it utilizes high-speed RoCE (RDMA over Converged Ethernet) for internal cluster communication. If there is an issue with Node 4's alignment within the GPnP profile or interconnect interfaces, it can experience internal delays that prevent the listener from processing incoming client requests seamlessly


On Exadata CC, the issue most frequently stems from resource saturation (reaching PROCESSES or SESSIONS limits), a connection storm causing registration lags between LREG/PMON and the SCAN listeners, or asymmetrical service states within the Guest VM Cluster (VMC). 

Step 1: Client and Architecture Log Frame Analysis
Before logging into the nodes, map out the connection path through the VMC architecture:
  1. Client Request: Hits the corporate DNS, which resolves the SCAN Hostname to 3 SCAN VIPs. 
  2. SCAN Listener: One of the 3 SCAN listeners receives the packet on port 1521 (Client network: bondeth0). 
  3. Redirection/Handshake: The SCAN listener looks at its service routing table to hand off the connection to a Local Listener on one of the 4 nodes via a dedicated database handler. 
  4. Failure State: If the database instance tells the listeners it is overwhelmed, the handler transitions to a state:blocked or state:ready with 0 capacity. The SCAN listener drops the connection and returns TNS-12516. 

Step 2: VMC Troubleshooting Steps & Detailed Commands
Log into the Exadata VMC as the grid user on any cluster node to diagnose the SCAN and Local listeners.
1. Check SCAN Listener Status and Handlers
Determine if the SCAN listeners are routing traffic properly or if the handlers are marked as BLOCKED. 
bash
# As grid user: Check the status of all 3 SCAN listeners
srvctl status scan_listener

# Check the active services and handler states on SCAN Listener 1
lsnrctl services LISTENER_SCAN1
  • What to look for: Look for your database service name. If you see handler: "DEDICATED" established:XXXX refused:XX state:blocked, the listener is intentionally rejecting connections because the underlying instance is reporting 100% utilization. 
2. Check Local VIP and Node Listener Status across the 4 Nodes
Ensure all 4 local node listeners are up and dynamically accepting registrations.
bash
# Verify clusterware status for all local listeners
srvctl status listener

# Check the local listener on the current node
lsnrctl status LISTENER
3. Parse the SCAN Listener Logs for Connection Storms
Cross-reference the timestamp of the client errors with the listener log file. On Exadata CC, Grid Infrastructure logs are located under the GI Oracle Base.
bash
# Navigate to the SCAN listener alert log directory
cd $ORACLE_BASE/diag/tnslsnr/<node_hostname>/listener_scan1/alert/

# View recent log entries for Refused or Error codes
tail -n 200 log.xml | grep -E "12516|12520|refused"
  • Interview Insight: If you see hundreds of connections incoming per second right before the TNS-12516 error spikes, you are dealing with an application connection storm. [1]

Step 3: Database Instance-Level Diagnosis (The Root Cause)
Switch to the oracle user and log in to the database via SQL*Plus to check if initialization limits have been reached.
1. Analyze Session and Process Utilization
Run this query across all 4 RAC instances to see if utilization is hitting 100% of the maximum allowed configurations. 
sql
SET LINESIZE 200
COLUMN RESOURCE_NAME FORMAT A15
SELECT INST_ID, RESOURCE_NAME, CURRENT_UTILIZATION, MAX_UTILIZATION, INITIAL_ALLOCATION 
FROM GV$RESOURCE_LIMIT 
WHERE RESOURCE_NAME IN ('processes', 'sessions');
  • Analysis: If CURRENT_UTILIZATION is equal to or dangerously close to INITIAL_ALLOCATION for processes, the LREG (Listener Registration) background process tells the listener to block new connections to protect the instance. 
2. Identify Connection Distribution per Node
See where the connections are stacking up. A misconfigured application profile might be bypassing the SCAN and overloading a single node.
sql
SELECT INST_ID, STATUS, COUNT(*) 
FROM GV$SESSION 
GROUP BY INST_ID, STATUS;
Step 4: Resolution & Remediations
Depending on your findings during the query steps, present these standard enterprise resolutions:
Scenario A: The Database Has Outgrown its Limits (Most Common)
If GV$RESOURCE_LIMIT shows processes are maxed out, increase the limits globally across the 4-node cluster. An Exadata X10M XL shape has massive CPU and memory capacity to handle higher boundaries safely. 
sql
# Increase limits in the SPFILE (Example: adjusting to 2000)
ALTER SYSTEM SET PROCESSES=2000 SCOPE=SPFILE SID='*';
ALTER SYSTEM SET SESSIONS=3000 SCOPE=SPFILE SID='*';
  • Note: Changing PROCESSES requires a rolling restart of the 4-node RAC instances to take effect. 
Scenario B: Mitigating Connection Storms & Listener Registration Lags
If PROCESSES are not maxed out but connections fail during traffic spikes, LREG might be failing to update the SCAN listener fast enough. 
  • Fix: Enforce application-side connection pooling (e.g., Oracle Universal Connection Pool) or set up Tailored Client Connect Timeouts and Retries in the client-side tnsnames.ora profile to handle minor registration delays seamlessly

Category 1: Oracle Exadata Core Capabilities
These foundational questions test your fundamental knowledge of Exadata's special sauce. 
Q1: What is Exadata Smart Scan and how does it improve query performance?
  • Answer: Smart Scan offloads query processing directly from the database compute nodes to the storage cell nodes.
  • How it works: Instead of sending whole database blocks across the network, the storage servers filter rows (predicate filtering) and isolate specific columns (column projection). Only the exact matching data is returned to the compute node, fundamentally eliminating network and database memory bottlenecks. 
Q2: What is an Exadata Storage Index and how is it managed?
  • Answer: A Storage Index is an in-memory metadata structure managed automatically by the storage cell servers to reduce physical I/O.
  • Key traits: It tracks the minimum and maximum values of table columns in 1MB chunks of physical storage. When a query executes with a WHERE clause, the storage server checks the index to skip reading 1MB blocks entirely if the data doesn't fall within the min/max range. It is never written to disk and requires zero manual upkeep. 
Q3: Explain Exadata Hybrid Columnar Compression (EHCC/HCC) and its distinct modes.
  • Answer: EHCC organizes data into logical compression units (CU) instead of traditional database blocks. Data within columns is grouped together, dramatically boosting compression ratios because identical column data types repeat. 
  • Modes:
    • Query High / Query Low: Optimized for analytic queries with fast decompression speeds.
    • Archive High / Archive Low: Optimized for maximum space reduction on cold data. 

Category 2: Exadata X10M Specific Upgrades
These questions test your familiarity with the latest architectural transformations introduced in the X10M platform. 
ASCII diagram

Q4: What is the most significant hardware shift in Exadata X10M compared to previous generations?
  • Answer: The complete transition to 4th Generation AMD EPYC (Genoa) processors across all database and storage servers. This migration provides up to 3x the core count in database servers (up to 96 cores per socket/192 per node) and doubles the processing power in the storage layer, directly facilitating up to 3x higher transaction throughput.
Q5: What replaced Intel Optane PMEM in Exadata X10M, and why?
  • Answer: Following Intel's discontinuation of Persistent Memory (PMEM), Oracle introduced Exadata RDMA Memory (XRMEM).
  • Technical details: XRMEM utilizes ultra-fast DDR5 DRAM inside the storage cells as a shared pool accessible via Remote Direct Memory Access (RDMA). It delivers incredibly low read latencies (<15 microseconds) for OLTP workloads without relying on physical PMEM hardware modules. [
Q6: How has the network infrastructure changed in the X10M architecture?
  • Answer: X10M introduces upgraded dual-port RDMA over Converged Ethernet (RoCE) network interface cards running on PCIe Gen 5. It delivers an active-active link configuration providing a total collective throughput of 200 Gbps per node. 

Category 3: Administration & Troubleshooting
These questions test your practical experience handling an Exadata environment. 
Q7: What utilities are used to administer Exadata Storage Cells, and how do they differ?
  • Answer:
    • CellCLI (Cell Command Line Interface): Used directly on a local storage cell to manage components like physical disks, cell disks, grid disks, and IORM profiles.
    • DCLI (Distributed Command Line Interface): A utility that allows a DBA to execute a single CellCLI command simultaneously across multiple or all storage cells in the rack. 
Q8: What is OEDA, and what is its role?
  • Answer: OEDA (Oracle Exadata Deployment Assistant) is a configuration utility used before physical installation. It captures customer infrastructure details (IP addresses, DNS, ASM disk layouts, hostnames) and outputs configuration XML/properties files required by the Oracle installation team to automatically provision the entire machine. 
Q9: What is the correct sequence to gracefully shut down an Exadata Database Machine?

  • Answer: To prevent data corruption, follow a strict top-down dependency sequence:
    1. Stop application services, database instances, and listeners.
    2. Stop the Oracle Clusterware stack (crsctl stop crs) across the compute nodes.
    3. Shut down the Database Compute nodes (shutdown -h now).
    4. Stop the storage services on the cells (alter cell shutdown all) and shut down the cell servers.
    5. Shut down RoCE/Infiniband and Cisco switches

Part 1: Top Oracle Exadata X10M Interview Questions & Answers
Q1: What makes the Exadata X10M architecture different from previous generations like the X9M?
  • Answer: Exadata X10M shifts its compute architecture by incorporating 4th Gen AMD EPYC processors, delivering up to a 3x increase in core density per database server (up to 190 usable cores per node). It also introduces PCIe Gen 5 routing, DDR5 DRAM memory, and replaces Intel Optane PMEM with Exadata RDMA Memory (XRMEM), maintaining an ultra-low 17-microsecond read latency. 
Q2: Explain how "Smart Scan" offloading optimizes query performance.
  • Answer: In traditional architectures, full database blocks must be transferred from storage over the network into the database server's SGA for processing. Smart Scan offloads query processing directly to the Exadata Storage Servers. The storage nodes handle row filtering (predicates) and column projection locally, sending only the requested data rows/columns back to the compute node. This eliminates significant network bottleneck and drastically cuts CPU cycles on the DB server. 
Q3: What is a Storage Index, and where is it kept?
  • Answer: A Storage Index is an in-memory structure maintained dynamically in the memory of the Exadata storage cells. It segments data into 1 MB storage regions and tracks the minimum and maximum values of specific columns inside that region. When a WHERE clause runs, Exadata evaluates the Storage Index first. If the target value falls outside the Min/Max range, it skips reading that 1 MB chunk entirely, bypassing physical I/O overhead completely. [
Q4: How does Exadata achieve low-latency communication between Database and Storage servers?
  • Answer: It uses the iDB (Intelligent Database) protocol mapped over a high-bandwidth 100 Gbps RoCE (RDMA over Converged Ethernet) internal network fabric. Because it leverages Remote Direct Memory Access (RDMA), compute nodes can directly fetch data blocks from the storage servers' XRMEM/Flash Cache without context switching or involving the storage OS kernel. 
Q5: Differentiate between the three Storage Server configurations available in X10M.
  • Answer:
    • Extreme Flash (EF): Contains all-flash NVMe drives designed for massive IOPS and high-throughput OLTP/Analytics.
    • High Capacity (HC): Balances cost and capacity by pairing high-density spinning hard disks (HDDs) with NVMe Flash Cards acting as a Smart Flash Cache tier.
    • Extended Storage (XT): Deep-archive tier with high-capacity disks but no flash acceleration, structured for low-cost historical data retention. 

Part 2: Architecture Comparison (Exadata vs. Excel Context)
Enterprise architects are sometimes asked how engineered solutions stack up against flat desktop workbooks when handling end-user reporting.
FeatureOracle Exadata X10M ArchitectureMicrosoft Excel Environment
Primary System UseEnterprise Mission-Critical OLTP / Big Data Data Warehousing.Desktop data analysis, local calculation, and presentation.
Data Capacity LimitsScalable into Petabytes (PB) per rack cluster via elastic expansion.Hard limit of 1,048,576 rows by 16,384 columns per worksheet.
Processing EngineDistributed scale-out architecture using 4th Gen AMD EPYC CPUs.Single workstation processing (constrained by local CPU/RAM).
I/O OptimizationSmart Scan, Storage Indexes, and Hybrid Columnar Compression (HCC).Loads entire dataset fully into local volatile memory.
Concurrency / AccessThousands of simultaneous transactional users via Oracle RAC.Typically limited to single or shared co-authoring file locks.



Question : exadatacc.x10m.xl architecture

exadatacc.x10m.xl refers to the Extra Large (XL) memory configuration shape for the Oracle Exadata Database Service on Cloud@Customer X10M infrastructure. 

This specific deployment model allows organizations to run highly performant, automated Oracle Databases and Oracle Autonomous Databases inside their own on-premises data centers while utilizing Oracle Cloud Infrastructure (OCI) management and billing. 
Core Hardware Specifications
The X10M-XL system shape scales out from a base configuration, offering massive processing power and the highest available memory footprint per database server in the X10M tier: 
  • Compute (Database Servers): Built on 4th Gen AMD EPYC processors. Each database server delivers 190 usable processor cores. 
  • Memory (Extra Large): Each database server features 2,800 GB (2.8 TB) of DDR5 RAM. This provides a massive capacity advantage for aggressive database consolidation and in-memory operations. 
  • Storage Tier: The initial base deployment starts with 3 Oracle Exadata storage servers, allocating 64 cores per storage server and featuring 1.25 TB of ultra-low latency Exadata RDMA Memory (XRMEM) per server to minimize read I/O bottlenecks. 
  • Network Fabric: Internal communication relies on a high-speed 100 Gbps RoCE (RDMA over Converged Ethernet) fabric. 
System Scaling Limitations
The infrastructure provides elastic expansion as resource demands change. Starting from the base layout, you can independently add compute or storage nodes up to the platform ceiling: [
  • Maximum DB Servers: Up to 32 nodes.
  • Maximum Storage Servers: Up to 64 nodes



Scenario 1: Storage Node Replacement & Graceful Maintenance
Question: You need to shut down an Exadata X10M Extreme Flash (XL) storage cell for physical maintenance. How do you ensure you will not cause an ASM disk group drop or database crash? What commands do you run?
Answer:
Before shutting down any Exadata storage cell, you must verify the ASM deactivation outcome. This checks if the remaining storage cells have enough redundant mirrors to maintain quorum and keep disk groups online. 

  • Step 1: Check if the cell can be safely taken offline. Run this command from the cell itself using CellCLI:
    bash
    CellCLI> LIST CELL DETAIL
    
    Look for the attribute asmDeactivationOutcome. It must explicitly say Yes. If it says "No", taking it down will drop your ASM disk group.
  • Step 2: Alternatively, run the check via Grid Infrastructure on a DB Node:
    bash
    $ORACLE_HOME/bin/kfod op=cellconfig
    

  • Step 3: Gracefully shut down the cell services: 
    bash
    CellCLI> ALTER CELL SHUTDOWN SERVICES ALL
    
Scenario 2: Smart Scan Not Working (Performance Drop)
Question: A batch query on an Exadata X10M database that normally runs in seconds is suddenly taking hours. You suspect Smart Scan Offloading is not engaging. How do you diagnose and verify this? 
Answer:
Smart Scan can fail to kick in due to factors like serial execution, wrong optimizer hints, or altered parameters. 
  • Step 1: Check session-level wait events in the database. Query V$SESSION_WAIT or look for specific Exadata offload wait events:
    sql
    SELECT event, total_waits FROM v$session_event WHERE sid = :sid AND event LIKE '%cell%';
    
    Look for cell smart table scan. If you see generic db file scattered read, Smart Scan is not working.
  • Step 2: Verify cell storage metrics directly. Execute via CellCLI to monitor whether the cell is actively offloading bytes:
    bash
    CellCLI> LIST METRICCURRENT WHERE name LIKE 'CL_BY_AND_REQ_W'
    

  • Step 3: Check Exadata Software Cell configuration. Ensure cell passthrough has not been forced:
    bash
    CellCLI> LIST CELL DETAIL
    
    Ensure cellPassthrough is set to FALSE. 

Scenario 3: Disk Degradation & Performance Profiling
Question: Users report intermittent I/O latency spikes on an X10M XL node. How do you isolate whether the root cause is a failing physical disk, a degraded Flash Cache, or a misconfigured IORM (I/O Resource Manager)? 
Answer:
You need to trace the metrics sequentially from physical layer to logical allocation layers using CellCLI. 
  • Step 1: Isolate physical or flash disk degradation:
    bash
    CellCLI> LIST PHYSICALDISK WHERE status != 'normal'
    CellCLI> LIST FLASHCACHE DETAIL
    

  • Step 2: Inspect individual flash disk or cell disk performance metrics:
    bash
    CellCLI> LIST METRICCURRENT WHERE name LIKE '.*_IO_RM_.*'
    

  • Step 3: Check I/O Resource Manager (IORM) status to ensure throttling isn't happening: 
    bash
    CellCLI> LIST IORMPLAN DETAIL
    
Scenario 4: DB Node to Storage Node Disconnection (RoCE Network)
Question: An Exadata X10M DB node cannot communicate with its storage cells. Given that X10M uses RoCE (RDMA over Converged Ethernet) rather than InfiniBand, what is your troubleshooting process? 
Answer:
On X10M, communication relies on RoCE Network Fabric via specific system configurations. 
  • Step 1: Validate the Compute Node cell configuration files. Make sure the DB node knows where to look:
    bash
    # cat /etc/oracle/cell/network-config/cellip.ora
    

  • Step 2: Use the Oracle Database tool to test if ASM can view the cells:
    bash
    $GRID_HOME/bin/kfod op=cellconfig
    

  • Step 3: If cells are missing from kfod, check RoCE link status using standard Linux/RoCE tooling: 
    bash
    # rdma link show
    # ibv_devinfo -v
    
Scenario 5: Overall Fleet Health Check Post-Patching
Question: You just finished patching an Exadata X10M environment. What tool and command do you run across the entire cluster to ensure no hardware or configuration anomalies remain? 
Answer:
You should use Oracle AHF (Autonomous Health Framework) / exachk. To sweep multiple cells or nodes concurrently, leverage the dcli (Distributed Command Line Interface) tool. 
  • Run a comprehensive cluster health check:
    bash
    # ahf check exachk -u -a
    

  • Check the hardware and firmware profile on all storage nodes simultaneously using dcli:
    bash
    # dcli -g cell_group -l root /opt/oracle.SupportTools/CheckHWnFWProfile
    
    A successful output will return [SUCCESS] The hardware and firmware profile matches... across all slots. 


System Profile & Architecture Highlights
  • Shape & Capacity: Exadata Cloud@Customer X10M.XL features high-core-count AMD EPYC (4th Gen) processors, DDR5 memory, PCIe Gen 5 NVMe flash, and redundant 100 Gbps RoCE (RDMA over Converged Ethernet) network fabric.
  • Virtualization Layer: Oracle Linux KVM (Dom0 hypervisor running KVM; user database cluster runs across DomU / VMC guests).
  • Database & Grid Infrastructure: 4-Node Oracle RAC running Grid Infrastructure 19c (RU 19.32) utilizing ASM/Flex ASM and Exadata Smart Scan offloading.
High-Yield Interview Q&A
Q1: "Connections to our 4-node RAC are hanging or experiencing intermittent connection timeouts. How do you diagnose and triage the SCAN and local listeners on ExaCC?"
Answer:
  1. Check VIP and SCAN listener distribution across the 4 nodes using srvctl status scan_listener and srvctl status scan.
  2. Determine if the hang is at the network handshake, SCAN redirection, or local listener process fork level.
  3. Check the listener log size and write rate; a bloated listener log or file system lock contention (log_file_size exceedance or logging hung on NFS/shared filesystems) will freeze incoming connections.
  4. Verify listener process responsiveness using lsnrctl status and strace to detect system-call blocks (e.g., futex, connect, or open blocks).
Q2: "What is a 'log frame' issue in an Exadata RAC context, and how does it manifest?"
Answer:
"Log frame" refers to latency or framing anomalies during Redo Log flush operations (log file sync / log file parallel write), RoCE packet framing drops (pause frames / PFC on RoCE v2), or buffer framing issues in the cluster interconnect (UDP/RDS over RoCE).
  • If redo write latency spikes, check whether FlashLog (Exadata Smart Flash Redo Log) is actively servicing redo writes on the cell nodes or if RoCE link drops/buffer pauses are occurring.
  • For interconnect/LMS log flushes, inspect GV$CLUSTER_INTERCONNECTS and dlv (data link layer) dropped frames via ethtool -S.
Q3: "Users report sudden query degradation across all nodes on an ExaCC X10M.XL cluster. Storage is normal, but CPU runqueues are spiking inside the DomU VM. How do you verify hypervisor (VMC/Dom0) CPU steal or oversubscription?"
Answer:
ExaCC runs on KVM hypervisors. Run top or mpstat 1 inside the DomU and inspect the %steal (%st) column.
  • If %st is significantly above 0.0%, the guest VM is ready to execute instructions but the physical host CPU (Dom0) is failing to allocate physical CPU cycles due to noisy neighbor VMs or CPU oversubscription.
  • Check Dom0 / KVM core pinning and virsh vCPU affinities using cell/compute management utilities or Oracle Cloud operations/SR.
Troubleshooting Steps & Diagnostic Commands
1. Listener & Network Connectivity Triage
Inspect both SCAN listeners and node-local VIP listeners across all 4 nodes:
bash
# Check status of Grid resources across all 4 nodes
crsctl stat res -t

# Check SCAN and Local Listener statuses via srvctl
srvctl status scan_listener
srvctl status listener -n <node1>

# Trace listener response time and check current endpoint queues
lsnrctl status LISTENER_SCAN1
lsnrctl status LISTENER

# Inspect listener alert/trace logs for ORA-12537, ORA-12541, or TNS-12535
# Path: $ORACLE_BASE/diag/tnslsnr/<hostname>/<listener_name>/trace/
tail -n 200 $ORACLE_BASE/diag/tnslsnr/$(hostname)/listener/alert/log.xml

# Check TCP socket queues (Look for non-zero Send-Q or Recv-Q on port 1521)
ss -tulpn | grep 1521
netstat -s | grep -i "listen drops"
2. Interconnect & RoCE "Log Frame" Diagnostics
Exadata X10M utilizes 100 Gbps RoCE network cards for RAC interconnect and Storage Cell communication:
bash
# Verify RAC interconnect interfaces and MTU (should be 9000 for jumbo frames)
ip link show | grep -E "re0|re1|bondeth"
ip -s link show

# Check for dropped frames, pause frames, and PFC buffer overflows on RoCE interfaces
ethtool -S <interface_name> | grep -E "drop|pause|discard|error"

# Check Cluster Interconnect performance from inside Oracle RAC
sqlplus / as sysdba <<EOF
SELECT inst_id, ip_address, is_public, if_name 
FROM gv\$cluster_interconnects;

SELECT inst_id, event, total_waits, time_waited_micro/1000000 time_waited_sec, average_wait 
FROM gv\$system_event 
WHERE event IN ('gc cr block lost', 'gc current block lost', 'log file sync', 'log file parallel write')
ORDER BY 1, 4 DESC;
EOF
3. DomU (VM Guest) & VMC (KVM/Dom0) Performance Troubleshooting
Isolate OS-level bottleneck from virtualization overhead:
bash
# Check CPU steal time (%steal indicates hypervisor throttling or noisy neighbors)
mpstat 1 5

# Inspect runqueue, context switches, and blocked processes
vmstat 1 10

# Verify HugePages allocation (crucial for 19c SGA on multi-terabyte X10M systems)
grep -i HugePages /proc/meminfo

# Check KVM hypervisor event metrics if hypervisor-level tools/exacli are accessible
exacli -e "list metriccurrent where name like '.*CPU.*'"
4. Exadata Storage & Cell Smart Log Performance
Confirm if Redo log flushes or offloading are being throttled at the Exadata Storage tier:
bash
# Check FlashLog throughput and latency from cell CLI (run on storage cell)
cellcli -e "LIST METRICCURRENT WHERE name LIKE 'FL_.*'"

# Check average cell read/write latency
cellcli -e "LIST METRICCURRENT WHERE name LIKE 'CD_IO_LOAD' OR name LIKE 'GD_IO_LOAD'"

# Run from database to check Exadata-specific wait events
sqlplus / as sysdba <<EOF
SELECT inst_id, event, time_waited_micro/1000/total_waits avg_ms
FROM gv\$system_event
WHERE event LIKE 'cell%' OR event LIKE 'log file%'
ORDER BY avg_ms DESC;
EOF


Scenario 1: Troubleshooting High Interconnect Latency & Packet Drops (RoCE Network)
Context: The X10M generation utilizes RoCE (RDMA over Converged Ethernet) via 100Gbps links instead of older InfiniBand. An interviewer might ask how you isolate severe performance degradation or "GC CR Block Lost" wait events across a 4-node RAC cluster.
  • Q: How do you check for packet loss or network issues on the RoCE interconnect fabric directly from the DB nodes?
    • A: Use the oracleorainfo or standard RoCE diagnostic tools such as ofed_info and ibpassthrough / roce_admin commands. However, the most definitive Exadata utility is exachk or tracking individual network interface counters via ethtool.
    • Command to run on the DB Node:
      bash
      # Check for drops, errors, or pauses on the RoCE interfaces (typically re0 / re1 or specific bonded interfaces)
      ethtool -S bondeth0 | grep -E "drop|error|missed|pause"
      
      # Check RoCE link status across all nodes using Autonomous Health Framework (AHF)
      ahfctl compliance --suite exachk
      
      Q: If Grid Infrastructure evicts a node due to communication failure on 19.32, what log files and commands do you use to pinpoint the exact root cause?
    • A: Check the clusterware alert log, ocssd.log, and cssdagent.log. Use crsctl to check cluster status, and analyze the CSS trace files.
    • Commands to run:
      bash
      # Check cluster status from any survival node
      crsctl check cluster -all
      crsctl stat res -t
      
      # View the summary of the node eviction using AHF/TFA
      tfactl diagcollect -crs -last 1h
      
      # Analyze the OCSSD log for network heartbeats
      view $ORACLE_BASE/diag/crs/$(hostname)/crs/trace/ocssd.log
      
Scenario 2: Smart Scan/Storage Offloading Tuning & Validation
Context: An application query runs slowly on your 4-node RAC. The interviewer wants to know if you can prove it is utilizing Smart Scan and leveraging the X10M's AMD EPYC optimized memory management unit (MMU) / XRMEM (Exadata RDMA Memory).
  • Q: How do you verify if a specific SQL ID is successfully executing a Smart Scan instead of pulling raw blocks over the RoCE network?
    • A: Query V$SQL_SYSTEM_STATS or V$SESSTAT for the session, or check V$SQL for columns tracking offloaded bytes.
    • SQL Commands:
      sql
      -- Check if a specific SQL ID is being offloaded
      SELECT sql_id, 
             io_cell_offload_eligible_bytes/1024/1024 AS offload_eligible_mb,
             io_cell_offload_returned_bytes/1024/1024 AS offload_returned_mb,
             (io_cell_offload_eligible_bytes - io_cell_offload_returned_bytes)/io_cell_offload_eligible_bytes*100 AS offload_efficiency
      FROM v$sql 
      WHERE sql_id = '&target_sql_id';
      
      Interview Tip: If offload_efficiency is close to 0%, the query is performing a block-by-block traditional read. Ensure cell_offload_processing = TRUE is enabled in your init.ora.
  • Q: How do you verify the health and status of the Flash Cache / XRMEM tier on the Exadata Storage Servers?
    • A: Log into one of the storage cells (using celladmin) and use cellcli.
    • Command to run:
      bash
      # Run from the cell server to see current caching status and health
      cellcli -e "LIST CELLSTAT LIST io_disk_stat_cache"
      cellcli -e "LIST FLASHCACHE DETAIL"
      
Scenario 3: ExaCC Automation & Cloud Agent Failures
Context: Because this is Cloud@Customer, infrastructure lifecycle tasks (like CPU scaling or local backups) rely on the Oracle Cloud Infrastructure (OCI) Control Plane agent (dbcsagent) running locally inside the DomU.
  • Q: During a high-load event on your 19.32 4-node RAC, online CPU scaling fails via the OCI console. How do you troubleshoot and fix the local agent framework?
    • A: Scale failures are often due to the dbcsagent hanging, running out of memory, or missing bindings in cprops.ini. You must check the status via dbaascli or systemd.
    • Commands to run:
      bash
      # Check the status of the local Cloud Control Plane agent
      systemctl status dbcsagent
      
      # View the agent log for API timeout or communication errors
      tail -n 200 /var/opt/oracle/log/dbcsagent/dbcsagent.log
      
      # Restart the agent if it is unresponsive
      systemctl restart dbcsagent
      
Scenario 4: Managing Node-Specific Asymmetric Performance (RAC Balancing)
Context: In a 4-node RAC cluster, one node often experiences much higher cell single block physical read times or high CPU utilization than the other three nodes.
  • Q: What RAC-specific dynamic performance views do you query to cross-compare Node 1 through Node 4 performance indicators in real time?
    • A: Use the global GV$ views to isolate anomalies across all 4 instances, specifically grouping by inst_id.
    • SQL Commands:
      sql
      -- Compare average wait times for crucial Exadata events across all 4 nodes
      SELECT inst_id, event, total_waits, average_wait
      FROM gv$system_event
      WHERE event IN ('cell single block physical read', 'cell smart table scan', 'gc cr block receive')
      ORDER BY event, inst_id;
      
      Q: If Node 3 shows massively inflated gc cr block receive times, how do you verify if the balance of the Grid Infrastructure Master services or LMS processes is bottlenecked?
    • A: Check the number and utilization of Global Cache Service (LMS) processes per node.
    • Command to run:
      bash
      # Check the OS level configuration of LMS processes
      ps -ef | grep lms
      
      # Use oradebug to check if LMS processes are being blocked or starved of CPU
      sqlplus / as sysdba
      ORADEBUG setmypid
      ORADEBUG lkdebug -m clustersvc
      
Scenario 5: Disk and Grid Infrastructure (ASM) Drop/Failures
Context: If a physical storage drive or grid disk reports errors, the 19.32 ASM instance should automatically drop or offline the disk depending on redundancy levels.
  • Q: A storage cell drops a grid disk. What ASM command verifies if rebalancing (ARBx processes) has automatically started, and how do you monitor its progress?
    • A: Query V$ASM_OPERATION to monitor active rebalance tasks across all 4 RAC instances.
    • SQL Commands:
      sql
      -- Check active ASM disk group operations and estimated time to completion
      SELECT group_number, operation, state, power, actual, so_far, est_minutes 
      FROM v$asm_operation;

Q1: How do you verify the status of all VIPs across a 4-node Exadata X10M cluster?
Answer:
You use the Oracle Clusterware control tool (crsctl) or Server Control utility (srvctl) from the Grid Infrastructure home. On Exadata X10M running 19.32, these utilities provide the real-time status and hosting node of each VIP.
Commands:
  • Check VIP status using srvctl:
    bash
    srvctl status vip -node exax10node1,exax10node2,exax10node3,exax10node4
    

  • Check cluster resource status filtered by VIP:
    bash
    crsctl stat res -t | grep -E "vip|ora.node"
    
Q2: Scenario: Node 3 fails, but its VIP does not failover to Node 1, 2, or 4. How do you troubleshoot this, and what Exadata-specific layer should you check?
Answer:
If a VIP fails to failover, it is usually due to network communication issues on the public network, incorrect subnets, or IP address conflicts (ARP caching issues). On Exadata X10M, the public network goes through high-speed RoCE (RDMA over Converged Ethernet) or traditional Ethernet depending on client network configuration, so you must verify the underlying network bond interfaces.
Troubleshooting Steps & Commands:
  1. Check Clusterware alerts: Inspect the Grid Infrastructure alert log (alert.log) for routing or binding errors.
    bash
    tail -n 200 $GRID_HOME/log/$HOSTNAME/alert.log
    

  2. Check IP allocation: Verify if another device on the network has stolen the VIP address using ping or arping.
    bash
    ping -c 3 <vip_ip_address>
    

  3. Check VIP resource configuration: Verify the target nodes and placement policies for the VIP.
    bash
    srvctl config vip -node exax10node3
    

  4. Force manual failover: If safe, manually migrate the VIP to verify clusterware responsiveness.
    bash
    srvctl relocate vip -node exax10node3 -targetnode exax10node1
    
Q3: How do you debug a VIP that is stuck in a STARTING or FAILED state on Node 1?
Answer:
A VIP stuck in STARTING or FAILED usually indicates that the Clusterware cannot plumb the virtual IP to the physical network interface (e.g., bondeth0). This can happen because the interface is down, the IP is already in use (duplicate IP), or there is an issue with the usm (Unified Storage Management) or network layer permissions.
Troubleshooting Steps & Commands:
  1. Check the exact resource error: Use crsctl to get detailed status text.
    bash
    crsctl stat res ora.exax10node1.vip -p | grep -E "STATE|LAST_SERVER|SCRIPT_TIMEOUT"
    

  2. Verify physical interface status: Check if the public network interface bond is up.
    bash
    ip link show bondeth0
    # Or check IP addresses currently assigned
    ip addr show
    

  3. Check for duplicate IPs on the network: Use arping from a different node to see if an MAC address responds to the failed VIP's IP.
    bash
    arping -I bondeth0 -c 3 <failed_vip_ip>
    

  4. Review the OS system log: Look for interface plumbing or network driver errors.
    bash
    journalctl -u crsd -n 100
    # Or traditional messages log
    tail -n 100 /var/log/messages | grep -i vip
    
Q4: In Oracle 19.32, what is the role of the Single Client Access Name (SCAN) VIPs compared to Node VIPs during a client connection phase on a 4-node cluster?
Answer:
The SCAN VIPs (typically 3 resolved via DNS/GNS) act as the initial point of contact for client connection requests. The SCAN listener routes the connection to the least-loaded node's local listener. The local listener then hands off the connection to the database instance via the Node VIP. If a node goes down, its Node VIP immediately fails over to an active node, allowing it to instantly send "Connection Refused" packets to clients attempting to connect to the dead node, preventing TCP timeout delays.
Commands to verify SCAN VIPs:
  • Check SCAN configuration and status:
    bash
    srvctl config scan
    srvctl status scan
    
Q5: How do you cleanly update or modify a Node VIP IP address in a 19.32 RAC environment?
Answer:
To modify a VIP address, you must stop the VIP resource, stop the dependent listeners, and then use srvctl modify nodeapps as the root user.
Sequence of Commands:
  1. Stop dependent database services and listeners on the node:
    bash
    srvctl stop listener -node exax10node1
    

  2. Stop the VIP resource:
    bash
    srvctl stop vip -node exax10node1
    

  3. Modify the VIP network attributes (Run as root):
    bash
    # Format: srvctl modify nodeapps -node <node> -address <vip_name_or_ip>/<netmask>/<interface>
    srvctl modify nodeapps -node exax10node1 -address ://example.com
    

  4. Start the VIP and listener back up:
    bash
    srvctl start vip -node exax10node1
    srvctl start listener -node exax10node1
Interview Question:
"In an Exadata Cloud@Customer X10M XL environment with a 4-node RAC running Oracle 19c (19.32), a node VIP fails to start or goes into an UNKNOWN state. Walk me through your step-by-step troubleshooting methodology and the exact commands you would use to isolate and resolve the issue." 

Answer: Structured Troubleshooting Framework
Step 1: Identify the Scope and State of the Cluster
First, check the global status of the cluster and pinpoint exactly which node's VIP is failing. 
  • Command: Run a comprehensive Clusterware resource check:
    bash
    crsctl stat res -t
    

  • Targeted Command: Isolate nodeapps and VIP status explicitly: 
    bash
    srvctl status nodeapps
    srvctl status vip -n <failed_node_name>
    

Step 2: Verify Network Interconnect & Public Interface Configurations
Exadata X10M environments use physical bonding for high-availability public networks. Ensure Clusterware's definition of the network subnet matches the OS configuration. 
  • Check Clusterware Network Config:
    bash
    oifcfg getif
    
    Ensure the public network interface (e.g., eth0 or bondeth0) is defined as public with the correct subnet/netmask.
  • Check Operating System Interface: 
    bash
    ip addr show
    # Or to view physical link state on Exadata:
    ip link show
    

Step 3: Rule Out IP Conflicts (ARP and Ping Checks)
A VIP will fail to start if another device on the network has inadvertently taken that IP address, causing an IP conflict.
  • Command: Attempt to ping the VIP from a remote node while the VIP resource is down:
    bash
    ping -c 3 <failed-node-vip-ip>
    
    If it responds while the resource is down, there is a duplicate IP conflict on the corporate network.
  • Clear ARP Cache (as root): If a VIP was recently relocated, clear adjacent switches' ARP tables or force an update: 
    bash
    arping -U -I <interface_name> <vip_ip>
    

Step 4: Analyze Cluster Logs via ADR and AHF
Since this environment runs Oracle 19.32, the Autonomous Health Framework (AHF) / tfactl is the fastest way to extract relevant logs. 
  • Command: Generate a targeted log collection for the network/VIP component over the past 2 hours:
    bash
    tfactl diagcollect -component crs -last 2h
    

  • Manual Log Inspection: Navigate to the Grid Infrastructure alert log to look for errors like CRS-5017 or CRS-2674:
    bash
    tail -n 200 $GRID_BASE/diag/crs/<hostname>/crs/trace/alert.log
    

Step 5: Attempt a Controlled VIP Relocation or Start
If the network is healthy and there are no IP conflicts, attempt to force the VIP back online or clear a frozen state. [
  • Start the VIP resource manually:
    bash
    srvctl start vip -n <node_name>
    

  • If stuck in a clearing/intermediate state, force relocation/restart (as root):
    bash
    crsctl stop resource ora.<node_name>.vip -f
    crsctl start resource ora.<node_name>.vip
    

Step 6: Exadata Cloud@Customer Management Layer Check
Because this is an Exadata Cloud@Customer platform, the problem could stem from the OCI Control Plane agent (dbcsagent) or hypervisor-level network security lists. 
  • Command: Check the health of the local Oracle Cloud automation agent:
    bash
    dbaascli agent status
    
    If the agent is hung, restart it to ensure grid/cloud state synchronization: 
    bash
    dbaascli agent restart


Scenario 1: RoCE Network Packet Drops & Interconnect Issues
Question: On an Exadata X10M 4-node RAC cluster running 19.32, you notice gc current block loss or gc cr block loss wait events spiking across the compute nodes. How do you troubleshoot the RoCE (RDMA over Converged Ethernet) network fabric to isolate the issue?
Answer:
Exadata X10M utilizes 100 Gbps RoCE instead of InfiniBand. High block loss indicates interconnect issues, packet drops, or RoCE configuration mismatches.
Troubleshooting Steps & Commands:
  1. Check the status of the RoCE network interfaces on the database nodes using ibv_devinfo or ip link.
bash
# Check the status and speed of the RoCE interfaces
ip link show
# Or use link layer tools
rdma link show
  1. Monitor RoCE port counters for dropped packets or errors using ethtool.
bash
# Check for Rx/Tx drops or pauses on the RoCE interface (e.g., eth8 or re0)
ethtool -S eth8 | grep -E "drop|error|pause"
  1. Use Exadata-specific tools like exadchk or sundediag to verify fabric health, or run ping with specific MTU sizes (Exadata uses MTU 9000 for RoCE).
bash
# Verify Jumbo Frames are working between compute nodes
ping -M do -s 8972 <interconnect_IP_node2>
Scenario 2: 19c RAC Node Eviction & Clusterware Troubleshooting
Question: Node 3 of your 4-node RAC cluster suddenly reboots (eviction). How do you investigate the cause of this eviction using 19c Grid Infrastructure logs and commands?
Answer:
Node evictions occur when a node fails to respond to network heartbeats (Network Heartbeat Failure) or disk heartbeats (Disk Heartbeat Failure), or if the OS hangs (Hang Manager).
Troubleshooting Steps & Commands:
  1. Use crsctl to check the current status of the cluster and locate which node is down or recovering.
bash
crsctl check cluster -all
crsctl stat res -t
  1. Run tfactl (Autonomous Health Framework) to automatically collect and analyze the logs from the time of the eviction.
bash
# Analyze the system for the past 2 hours to find the root cause
tfactl analyze -last 2h
  1. Manually check the Cluster Health Monitor (CHM) and the Grid Infrastructure alert log found in the GI App Diagnostic Dest.
bash
# View the alert log for Grid Infrastructure on the evicted node
tail -n 500 $GRID_HOME/log/<hostname>/alert<hostname>.log
  1. Examine the CSSD (Cluster Synchronization Services) daemon log to see if it was a network or disk timeout.
bash
tail -n 1000 $GRID_HOME/log/<hostname>/cssd/ocssd.log | grep -E "clssnm|eviction|timeout"
Scenario 3: Smart Scan (Cell Offload) Failure Post-19.32 Patching
Question: After applying the 19.32 Release Update, a critical batch job slowed down. Execution plans show that the queries are performing standard sequential scans instead of Exadata Smart Scans. How do you troubleshoot why Cell Offload is not working?
Answer:
Smart Scans can fail or be disabled due to mismatched parameters, cell/compute software version compatibility issues, un-indexed or un-aligned data blocks, or if the initialization parameters are altered.
Troubleshooting Steps & Commands:
  1. Verify that Exadata cell offload features are globally enabled in the database instance.
sql
SHOW PARAMETER cell_offload_processing;
-- Ensure it is set to TRUE. If not:
ALTER SYSTEM SET cell_offload_processing=TRUE SCOPE=BOTH;
  1. Verify that the compute nodes can communicate with the Exadata Storage Servers using cell_cells.ora.
bash
cat $ORACLE_HOME/dbs/cell_cells.ora
  1. Check cell server status and statistics from the CellCLI utility on one of the storage cells to ensure offloading isn't disabled due to high cell CPU/memory load.
bash
# Run on the storage cell
cellcli -e "LIST CELL ATTRIBUTES bmsConnections, status"
  1. Trace the session executing the query to see why offload was disabled using the cell_smart_scan_failover session statistic.
sql
SELECT name, value FROM v$mystat m, v$statname s 
WHERE m.statistic# = s.statistic# 
AND name LIKE 'cell%smart scan%';
Scenario 4: High ASM Rebalance Times on Exadata X10M Storage
Question: A flash drive failed on one of the X10M Storage Servers, triggering an automatic ASM rebalance. The rebalance is taking too long, impacting OLTP performance on the 4-node RAC. How do you check the progress and safely optimize the rebalance speed?
Answer:
Exadata X10M features extreme high-speed NVMe storage. ASM rebalances should be fast, but if the ASM_POWER_LIMIT is set too low, it will bottleneck. If set too high, it might conflict with database I/O.
Troubleshooting Steps & Commands:
  1. Connect to the ASM instance (+ASM1) and check the status, estimated time, and power of the ongoing rebalance operation.
sql
SELECT group_number, operation, state, power, est_minutes FROM v$asm_operation;
  1. Check the performance impact and throughput of the rebalance using v$asm_diskstat.
sql
SELECT name, reads, writes, read_time, write_time FROM v$asm_diskstat WHERE group_number = <group_id>;
  1. Dynamically adjust the ASM power limit to speed up the rebalance. In 19c, the power limit can scale up to 1024.
sql
-- Increase power limit to accelerate execution (tune based on system load)
ALTER DISKGROUP <diskgroup_name> REBALANCE POWER 64;
Scenario 5: Managing Rogue Sessions causing High CPU on X10M Compute Nodes
Question: Compute Node 1 is experiencing 98% CPU utilization due to multiple parallel query sessions spawned from a 19.32 database instance, threatening the stability of the other instances. How do you quickly identify and safely terminate these rogue processes?
Answer:
Using a combination of OS commands and SQL queries, you can trace parallel execution coordinator sessions down to their slave processes.
Troubleshooting Steps & Commands:
  1. Run top or htop on the affected node to find the top consuming OS process IDs (PIDs).
  2. Query v$session and v$px_session to cross-reference the OS PID with the database session and see what SQL text is executing.
sql
SELECT q.sql_text, s.sid, s.serial#, s.osuser, s.machine, p.spid 
FROM v$session s 
JOIN v$process p ON s.paddr = p.addr
JOIN v$sql q ON s.sql_id = q.sql_id
WHERE p.spid = '&OS_PID';
  1. Kill the coordinator session gracefully inside the database to cleanly terminate all associated parallel slave processes across the 4 nodes.
sql
ALTER SYSTEM KILL SESSION 'sid,serial#' IMMEDIATE;


Troubleshooting an Exadata Cloud at Customer X10M XL (exadatacc.x10m.xl) platform involves a multi-tiered approach spanning the VM guest compute tier, the Exadata Storage servers, and the central control plane. By combining granular CLI tools (ExaCLI, dbaascli, dcli) with the AI-driven automation of Oracle Enterprise Manager 24ai, you can rapidly isolate whether an issue stems from database resource contention, storage cell misconfigurations, or network dropouts. 

Step-by-Step Troubleshooting Flow
Step 1: High-Level Diagnostics via Enterprise Manager 24ai
Before logging into individual servers, use the centralized console to locate the root cause: [
  • Use the GenAI Assistant: In the EM 24ai console, type natural language commands like "Show me top 3 databases by I/O wait on my ExaCC cluster" or "Analyze storage cell alerts". The assistant will automatically generate real-time performance dashboards. 
  • Navigate to Exadata Target Infrastructure:
    1. Click Targets → All Targets → Select your Oracle Exadata Database Machine target.
    2. Inspect the Exadata Storage Server topology maps to check for cells with red/yellow thresholds.
    3. Under Performance, check the IORM (I/O Resource Management) tab to confirm whether resource plans are throttling lower-priority workloads. 
Step 2: Collect Cloud Tooling & Agent Health Status (CLI)
If EM 24ai reports missing data collections or an unreachable target status, verify the local cloud agent frameworks on the VM Guest: 
  • Check the Cloud tooling agent status:
    bash
    # Check the cloud management tooling agent status on the VM
    sudo systemctl status dbcsagent
    

  • Force collection refresh via Management Agent:
    bash
    # Force EM agent to re-collect the Exadata configuration metrics
    emctl control agent runCollection <target_name>:oracle_exadata <collectionName>
    
    Step 3: Storage-Tier Diagnostics with ExaCLI
For an exadatacc.x10m.xl deployment, traditional CellCLI is restricted; secure administration is performed using ExaCLI from the VM guest or client endpoint. 
  • Verify ExaCLI credentials and list storage cell status:
    bash
    exacli -c celladmin@<storage_cell_IP> -e "list cell attributes name, status, cellVersion"
    

  • Diagnose active alerts and disk issues on the cell:
    bash
    exacli -c celladmin@<storage_cell_IP> -e "list alerthistory where alertType='State' and severity='Critical'"
    

  • Examine Flash Cache and Storage Index Efficiency:
    bash
    # Ensure the massive X10M Flash cache capacity is operating efficiently
    exacli -c celladmin@<storage_cell_IP> -e "list metriccurrent where name like 'CL_BY_AND_REQ_F'"
    
    Step 4: Executing Cluster-Wide Health Assessments
When multiple cells or VM hosts are involved, leverage distributed commands and Autonomous Health Framework (AHF).
  • Check storage configuration mapping from the compute node:
    bash
    $ORACLE_HOME/bin/kfod op=cellconfig
    

  • Run a cluster-wide storage alert check using dcli (if accessible):
    bash
    dcli -g cell_group -l celladmin "cellcli -e list alerthistory"
    
Summary of Troubleshooting Tools
Tool / InterfaceBest Used ForSample Action
Enterprise Manager 24aiAI-driven target discovery, global IORM visualization, metric visual graphs.Ask GenAI Assistant: "Show performance anomalies".
ExaCLIRemote cell monitoring, physical disk and Exadata Flash tier troubleshooting.exacli -e "list physicaldisk"
dbaascliDatabase lifecycle execution logs, primary/standby patching failures.dbaascli database status
emctlTroubleshooting blocked management agents or fixing collection lag.emctl status agent