Streamline Your Infrastructure Management and IT Operations
Recently active
When it comes to passwords, I am the worst. I can not remember a single password I have. Worse, I struggle to type it correctly. In some way, this is an advantage as I have to reset all my passwords frequently. While you, reader, may think that this does not apply to you, I am confident I am not alone in my struggle. Incidentally, I have seen a few cases dedicated to being unable to log in to Prism with admin credentials or to a CVM as nutanix. Below is a very generic approach on how to handle the issue.First of all, I feel compelled to remind you that there is a lock out timeout for consecutive failed authentications. By default, the timeout is 15 minutes, but you can unlock the account from a CVM when logged as nutanix.In some cases, a password is simply unknown for whatever reason. Again, if you know nutanix user credentials, you can log in to CVM and change the password.What if you are certain that the password is correct and, maybe, even unlocked it already but the Prism still doe
If you have decommissioned a vCenter without unregistering it with the Prism and now are unable to complete the task, there may be a solution.For AOS 5.5.x provided the symptoms match an upgrade may be required to resolve the issue. If you see all of the following, upgrade to AOS to 5.5.6, 5.8.2, 5.9 or later.An attempt to unregister a vCenter instance that does not exist anymore from Prism using vCenter Registration page creates a task that fails at 100% progress. Using the "ncli managementserver unregister" command produces the same outcome. Uhura logs (home/nutanix/data/logs/uhura.out) show vCenter is not reachable and is not able to delete the connection object.Nothing is fully predictable - we understand that. Some things still can be planned. We encourage you to plan the maintenance, especially when it comes to decommissioning of the components. It is generally more difficult to solve the problem where a component no longer exists then while it was still available.For more detail
To provide access to cluster storage, Nutanix Volumes utilizes an iSCSI data services IP address to clients for target discovery. This iSCSI data services IP address acts as an iSCSI target discovery portal and initial connection point. Nutanix does not recommend configuring iSCSI client sessions to connect directly to Controller VM IP addresses. Changing Data Services IP address requires downtime and planned maintenance for affected services and applications. Prior to changing the Data Services IP address, check if any of the features below are enabled on the cluster, such as Calm, Leap or Xi Leap, Karbon, Objects or Files. If none of the features is enabled, proceed with changing the Data Services IP Address as per the Prism Web Console Guide: iSCSI Data Services IP Address Impact.Karbon and Objects. If Karbon and/or Objects Services are enabled do NOT change the Data Services IP address. Reach out to Nutanix Support for assistance.Nutanix Files. If Data Services IP is reachable and
The Health Page within Prism is a great resource to keep track of the status of your environment. With a simple dashboard layout, the health page provides checks for many components of the Nutanix environment. This dashboard is dynamically updated with health information based upon Nutanix Cluster Checks (NCC) which execute at specific time intervals.Historical data is also captured and presented in this dashboard using a table and graphical layout. For example, if a check has reported as failed, reviewing the historical data can help determine if the failure was the result of a change within the environment at a particular time. This data also assists Support in determining where to review in logs if further troubleshooting is needed.For more information about the Health Dashboard, please review the associated documentation found in the Support Portal.
It’s over a hump day so a quick tip today.In AHV if you ever attach a VM’s disk to a Windows VM and then delete the disk owner VM you will still see the disk as attached and browsable. Refreshing disk information in Disk Management utility in Windows will result in marking the disk as “Not Initialised”. Rescan the disks instead.For more details refer to KB-7368 AHV | Deleted VM disk attached to Windows VM is accessible from guest OS
I am confident that by now you have heard of Xi Leap and Leap. While both of the products have the same purpose they are quite different. Xi Leap will recover into the Cloud. With Leap, however, you recover to another of your sites.Leap offers an entity-centric automated approach to protect and recover applications if there is a failure. It uses categories to group the VMs and automate the protection of the VMs as the application scales. Application recovery is more flexible with network mappings, an enforceable VM power-on sequence, and inter-stage delays. Application recovery can also be validated and tested without affecting your production workloads. Asynchronous, NearSync, and Synchronous replication schedules ensure that an application and its configuration details synchronise to the recovery location for a smoother recovery.Roughly, what are the options here? You can recover within Prism Central management domain i.e. between clusters or sites that are managed by the same PC. Yo
If you are running AOS 5.10.9 or below or AOS 5.11.1 and thinking of upgrading or have already scheduled one there is on thing to consider prior to executing the task.Check the size of localhost_access_log.txt file. If it has grown out of proportion (several GB) then you may want to clear it first and, if you can, chose AOS 5.10.10 or later, AOS 5.16.1 or AOS 5.17 or later as an upgrade target.Otherwise you may find some issues with the upgrade.Save yourself some hassle and see the details in the KB-8561 Prism Central/Prism Element upgrade failing with /home space constraints due to localhost_access_log.txt files not rotating
AOS, LCM and NCC Versions We can check the versions from the about section on the Prism GUI. Log in to the Prism UI. Click the drop-down next to your username in the top right corner. Click the "About" button. Here you will find the AOS, LCM and NCC versions. AOS, LCM, NCC, Foundation, Files and AHV AOS, LCM (Life Cycle Manager), NCC, Foundation and AHV versions can be found in LCM Inventory Page. Refer to the relevant LCM Guide in Nutanix Portal here to find the latest instructions. Below are steps to find them from AOS 5.10.x. Log in to the Prism Element UI. Click the gear icon (Settings) in the top right and select "Life Cycle Management". Click Perform Inventory. Once the inventory is complete, the same page will display the versions of the different components ERAAs soon as you log in to the Era UI Home Page, there will be a panel with the Era version listed.MoveUse below steps to find the Move version: Log in to the Move UI. Click your username in the top right an
When maintaining a Nutanix infrastructure, there is sometimes a need to mount an ISO to a host. For example, when performing a SATADOM replacement or booting into Phoenix, an ISO needs to be mounted via the IPMI interface.When using the IPMI GUI, you will likely be attaching an ISO from your local computer which is most likely outside the CVM/Host’s network. Therefore, the IPMI/hosts ability to read from the attached ISO can be hampered by network connectivity issues. This is especially true with many workers currently working from their homes.With the SMCIPMITOOL found on the CVM, you can leverage another CVM in order to mount an ISO. To note, the ISO being mounted would first need to be copied to the local filesystem of the sharing CVM (the path /home/nutanix/tmp could work to store the ISO - just be mindful of the available free space of the /home partition).More information, including procedural steps, can be found in KB 4616
With automatically generated alert cases, Nutanix will create cases based upon certain alerts and proactively reach out to customers to troubleshoot and resolve these alerts. Accordingly, ensuring that the primary support contact details are correct helps streamline the support process.When a company has personnel that have departed or new personnel are placed in-charge of a specific Nutanix asset, it is important to keep the contact details for the asset up-to-date to ensure that a support agent makes contact with proper personnel.NOTE: The primary support contact is also the default contact for shipping dispatches, although this contact information is re-confirmed throughout the dispatch process.More information about contact details and the procedure for changing them can be found within the Support Contact Details knowledge base article (KB 2307).
Is it really possible that a networking issue, which exists on the other side of a large/vast network, could manifest locally on a host as NIC CRC errors (rx_crc_errors)?Yes!The way that frames are moved across a network with cut-through switching (which is the model used by current/high-performant data center switches) differs from the traditional store-and-forward model.With store-and-forward switching, frames are entirely received (and error-checked) on a switch before being passed along to the next switch inline. If errors are found within a frame, the frame is not passed along to the next switch.However, with cut-through switching, a switch is simultaneously receiving a frame and already passing it along to the next switch inline! Error checking is then only completed on a frame after it has already left a switch. Accordingly, any errors within the frame simply get passed along to the next switch inline until the frame reaches its final destination.In this way, when troubleshooti
Sometimes, through the normal operation of a hypervisor host, network interfaces (NICs) can become “overwhelmed” and be unable to respond to traffic fast enough as being served by an upstream network switch. This condition results in frames, which are subsequently being dropped from the receive buffer of the host, to go unprocessed or “missed”.Though it is generally not a good situation when a host (and/or a downstream element from the host, like a VM) “misses” network traffic, many applications are robust enough to handle a few misses from time to time. However, if this condition is frequent and/or perpetual, it can cause production issues and alarms would be exhibited from Prism accordingly.To combat a frequent/perpetual condition, there are several available options. For ESX and Hyper-V hosts, the receive buffer size of NICs can simply be increased. For AHV hosts, instead of increasing the receive buffer size, it is recommended to employ load-balancing across the available uplinks.
Hi,I'm looking up for a consolidated monitoring sheet if any available on Nutanix AOS and ESXI hyper-v integrated with some threshold applicable and severity will be really helpful. If someone has please share it.
When you are adding new features to your Nutanix environment, there might be a need to update the license or purchase a new one. If you do not update the license, there will be a red banner at the top of the Prism Page indicating that there are licensing violations.Clicking on “View Licensing Details” will show offenders of the current license. Using NCLI, you can see more details about features available in the current license level and features that are not allowed.For more information about disabling features, attaining the correct license, or temporarily removing the banner, look at KB 3443.
Question: Does nutanix use Self Signed Cert or CA-cert by default?
This article’s purpose is to provide assistance in determining the right amount of resources required to enable various services on Prism Central.Currently, there are multiple services available such as CALM, Karbon, Objects, Files, etc. Each of these services adds their own resource requirements to Prism Central hence, increasing the need for more resources for Prism Central in order to function properly. If enough resources are not configured, you might encounter the alert “Configured resource for the Prism Central VM is inadequate”. The NCC health check pc_vm_resource_resize_check verifies that the configured amount of memory and vCPU resources of the Prism Central VM is adequate. This check is introduced in NCC release version 3.10.0 and applies to Prism Central VMs (version 5.17.1 or higher) only (inapplicable to Controller VMs (CVMs) of Prism Elements clusters).Please refer to KB-8932 for more details.To learn more about the various services available and to determine the Prism C
Let’s face it, upgrades are daunting, and confusing, and frequent, and unavoidable, and, well, painful. Aiming to help with the preparation process and alleviate at least some of the worries Nutanix put together Acropolis upgrade (Upgrading AOS, Prism Central, Hypervisors, and Related Software Through The Web Console).Recommended Upgrade Order:Prism Central (PC): Upgrade and run NCC on Prism Central. PC: Upgrade Prism Central. PC: Run NCC. Prism Element clusters (PE): Upgrade and run NCC. PE: Upgrade Foundation. PE: Run and upgrade Life Cycle Manager (LCM): Perform an LCM inventory (also updates LCM framework). Do not upgrade any other software component except LCM in this step. PE: Upgrade AOS. PE: Run and upgrade Life Cycle Manager (LCM): Perform an LCM inventory (also updates LCM framework). Upgrade SATA DOM firmware (for hardware using SATA DOMs) as recommended by LCM. Upgrade all other firmware as recommended by LCM (BIOS / BMC / other). PE: Upgrade AHV for AHV clusters. PE: U
Have you been trying to keep your virtual machines (VMs) up to date with the latest versions of Nutanix Guest Tools (NGT), but face following hurdles:More difficult with seemingly ever-increasing release version activity? Tired of seeing alerts regarding a new version of NGT being available for a VM from within the Prism?Well, the good news is Prism Central is the answer for all of your NGT version management needs. Within Prism Central, NGT management has been greatly improved and simplified.For example, you can select particular VMs that are to receive an upgrade and employ this upgrade all at once (rather than one at a time and manually as through Prism Elements). Further, you can select the options to initiate the upgrade(s) immediately or, later, at scheduled time of your convenience.You can find more information about NGT management with Prism Central as per the “NGT MANAGEMENT IN PRISM CENTRAL” section of the Prism Central Guide.
First off, Prism Central RBAC SUCKS compared to vSphere. 11 years supporting VMware, which I am completely self taught, the the simple task of creating a role like the VMware 'VM Power User’ role has me ready to update my resume. This is maddening. I need a role that will be assigned to team members to manage all aspects of VMs except creation/deletion, and they need to be able to mount ISO images to the VMs. I have gone through the granular list of permissions, and I don’t see anything like that under VM or Images. Where is this hidden?
This article aims to better explain the nuances of the alert “Disk space usage for one or more disks on controller VM” as seen in Prism and I have included some basic troubleshooting steps to identify the exact issue. SSH into any CVM of the cluster and run the below command: ++ allssh “df -h”This will list the use % of all the disks controlled by the CVM. For ex: see screenshot belowNutanix reserves space in the SSD under /home for its infrastructure and is capped at 40GB, and sometimes it is possible to run low on the space usage and triggering an alert. Check out this article for more information on how to reduce /home space usage. The other disks /dev/sdx are the individual physical drives on each node. First we eliminate the possibility of a faulty drive by running a smartctl check. ++ sudo smartctl -a /dev/sdx Replace x with the appropriate disk name from df -h output. If the smartctl test fails, then contact your vendor for disk replacement. Back to the df -h output, the s
Sometimes Kubernetes clusters are not listed when we try to open Karbon console from Karbon management GUI on Prism Central. This is due to Karbon (karbon_core & karbon_ui) containers being in an unhealthy state. This can be due to various reasons, one of them being docker containers are unable to get the proper resources or docker is unable to establish a connection with Karbon containers. Below is a screenshot of what one would notice when one would go to the Karbon console and no Kubernetes clusters are listed. Symptoms are container is unhealthy and here is the simple fix to try as a first step: The Karbon container karbon_core is unhealthy. The Karbon container karbon_core is unhealthy.nutanix@NTNX-10-X-X-X-A-PCVM:~$ docker ps CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES 3cf0cb9a9dee karbon-core:v1.0.1 "/start.sh" 5 months ago Up 5 months (unhealthy)
How can we determine if Node is experiencing NIC issues and if so, then what? There are many reasons that can cause host NIC errors and troubleshooting this usually involves analysis to triangulate the issue to a certain part in the networking topology. Link flapping (interface continually goes up and down) Cable disconnect/connect Faulty external switch port Misconfiguration of the external switch port Faulty NIC port Faulty cable Faulty SFP+ module The two big error counters that we are concerned are rx_crc_errors and rx_over_errors (and in conjunction, rx_missed_errors/rx_fifo_errors).Run the following command on your host depending on the hypervisor:AHV:ethtool -S <eth> | egrep "rx_errors|rx_crc_errors|rx_missed_errors"ESXi:esxcli network nic stats get -n <vmnic> | egrep "Total receive errors|Receive CRC errors|Receive missed errors"Hyper-V:Get-NetAdapterStatistics -Name Ethernet*<interface number> | fl * Rx_crc_error:The sending host computes a cyclic
Anecdote by @Kiran Surya Nutanix users and administrator are often required to access the CVMs, Hypervisors and IPMIs through their IP addresses. The most commonly used method to look at the IP addresses of these three entities is through Prism Element, and going through the Hardware tabs. However, looking through the GUI can be time consuming since you are required to click on every Host to view the IP address details of all three components. It’s all the more cumbersome if you have a huge infrastructure, like say 24 Hosts. It would be ideal if there is an easier process to list all these pieces of vital information. I happen to use the following combination of characters to list all these three pieces of data, and needless to say the customers are delighted when they see me run this command: nutanix@NTNX-7B43JB2-A-CVM:10.171.154.127:~$ nutanix@NTNX-7B43JB2-A-CVM:10.171.154.127:~$ echo; set `ncli host list | grep Address | sed 's/^.*: //'`;echo -e " IPMI\t CVM\t\tHy
Original plans do not always work out as intended. Goals change over time, new features are introduced. The modern world is an ever-changing one. You decide to migrate Prism Central VM (or even all three of them in case of a Scale Out PC deployment) to a different container but alas see no ‘migrate’ button. You can still do it.Find VM disk files Power off the VM/VMs Create a new Image for each VM disk Create a new VM/VMs using acli or using Prism UI. Use image created in Step 3 as a source for the disk. Power on new VM/VMs and check functionality. If static IP address was used you may need to manually configure it one more time. Delete old VM if it is not needed anymore Delete Image if it is not needed anymoreThat’s right, these are the same steps used for migration of a regular (User VM) between containers.For a detailed process with commands and screenshots refer to KB-2663 AHV | How to move VMs between containers within an AHV cluster
Are there any recommendation for protecting Prism Central - I’m actually using PC Scaleout so have 3 PCVM’s and am curious if I should place these in a Metro or Async Protection Domain in the event that we perform a datacenter failover.I know that this cluster wouldn’t be critically needed in such a scenario but it would be nice to have. Any pro’s/con’s to using either PD type (Metro vs Async) for this?
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.