The Cause of the Mismatch
The InstanceNotFound error occurring while an instance still appears in openstack server list is typically an expected behavior of the OpenStack asynchronous deletion lifecycle. When a delete command is issued, Nova marks the instance for deletion in the database and triggers a series of teardown tasks across the Nova Compute and Conductor services. The instance remains in the list because its database record exists until the final cleanup task completes.
The error occurs because openstack server show attempts to retrieve a full object detail that may no longer be available to the API's active lookup logic, or the record has transitioned to a state where it is logically "gone" but not yet physically purged from the database.
Cache Consistency vs. Performance
The oslo.cache expiration is a trade-off designed to protect the Nova database from excessive read pressure in large-scale deployments. In most production environments, relying on manual cache flushes is not recommended as it introduces operational overhead and does not address the underlying asynchronous nature of the Nova state machine.
Automatic invalidation is feasible, but the "stale" view is usually a symptom of the asynchronous deletion process rather than just a cache timeout. Even with a zero-second cache, the instance may appear in the list until the Compute node confirms the VM is destroyed and the Conductor removes the DB entry.
Recommended Approach
- Accept Asynchronous Latency: Treat the
DELETING status as the source of truth. If the instance is visible but returns InstanceNotFound on a detail query, the deletion is already in progress.
- Avoid Aggressive Cache Tuning: Keep
oslo.cache at default values unless you observe significant API latency. Reducing this too far can lead to database contention during peak load.
- Verify Actual State: Use the following command to check if the instance is stuck in a specific state:
openstack server show <instance_id>
If the command fails with InstanceNotFound but the list persists, monitor the nova-compute logs on the host where the VM resided to ensure the teardown is not hanging.
Diagnostic Requirement
To determine if this is a cache issue or a stuck deletion process, please provide the vm_state and vm_status from the Nova database for the affected UUID. If the state is deleted but the record persists, it is a database cleanup failure; if it is deleting, it is standard asynchronous behavior.