Streamline Your Infrastructure Management and IT Operations
Recently active
Q.Which should be upgraded first BIOS, BMC or Host Boot Firmware?A. Please find below the upgrade order for all Nutanix firmware and software components. For hardware that has SATADOM, SATADOM should be upgraded first before the rest of the components. Always update BCM before the BIOS.Acropolis Upgrade Guide v5.17 Q. Can numerous upgrades be performed at the same time? For example, can I upgrade Boot Firmware on all five hosts in the cluster at once?A. You can select all components for all hosts to be upgraded. However, the upgrade will be done on each component of a host one by one. So only one component of one host will be upgraded at a time. When completed the LCM will proceed with the upgrade of the next component and so on. The non-critical and non-impacting components are upgraded first. Q. Should these upgrades be performed after hours? Is there an estimated timeframe for how long these upgrades take?A. No, the LCM upgrades are non-disruptive, so it does not require downtime.Th
Keeping track of hardware details across a cluster can be overwhelming. Some information could be found in Portal and other information might be found in IPMI.NCC provides a tool via the command line that allows for the accumulation of all the necessary details into a single output using the following command: ncc hardware_info show_hardware_info.You can use this command to see BIOS/BMC versions, DIMM Serial numbers, Disk firmware versions, etc. More information about the data collected can be found in KB 7084.This information is stored in the CVM for that specific node for access later and is collected by the sysstats utility every 24 hours. More information about the location can be found in the Support Portal Documentation.
Prism Central has Image management feature that allows pushing disk images to managed Prism Element instances based on configured image placement policies. The policies rely on categories that are associated with both images and clusters. Altogether this allows building a granular and intricate (if you chose to) set of image distribution policies.If for any reason using Image Management is not an option for you, you can push images from the PC to the clusters of choice manually either using PC CLI. Read further for details and instructions:Prism Central Guide v5.17: Image ManagementKB-4892 Prism Central: Manually pushing images uploaded to PC to multiple Prism ElementsKB-6205 Prism Central - How to create VMs on containers other than SelfServiceContainer
Would you like to never see a PSOD screen again? I mean, who would? Elevated heart rate, restarted VMs, explaining the reason, searching for the reason, trying to distinguish a false positive from an actual alert. You definitely have better things to do.When looking at memory failures (DIMM is what I am referring to in this instance) uncorrectable memory errors is what brings most of the trouble and requires DIMM replacement but the majority of the errors are correctable.Each BIOS version brings improvement to handling hardware. That includes correcting memory errors. Meaning, you can avoid pesky hardware replacement windows for much longer if not forever.To keep your BIOS up-to-date use LCM. If you do opt for a manual process remember to always upgrade BMC first and only then proceed with the BIOS. BMC upgrades are non-disruptive. BIOS upgrade does require a host reboot.And as always, keep your NCC up to date (current NCC version is 3.10.0.1 and some false positive alerts in relation
I thought I’d put together a set of sources to help with troubleshooting some of the popular issues with the Prism.If you ever lock yourself out of Prism account do not fret. Local admin account can be unlocked by logging into a CVM with nutanix user and resetting the password for the admin user.KB-8130 Prism Admin Account Locked With "Account locked due to too many failed attempts." If you ever need to find a Prsim Leader node as part of any troubleshooting:KB-1841 Checking which node is the Prism Leader. To verfify connectivity status between Prism Central and Prism Element use NCC check cluster_connectivity statusKB-3379 NCC Health Check: cluster_connectivity_status. For troubleshootimg PE to PC connectivity follow KB-6970 PE-PC Connection Failure alerts.Expand the list in the comments, help the community!
Did you know you can determine when a VM was created? All you need is a Prism Central instance.In PC UI go to Virtual Infrastructure –> VMs –> List –> in Focus dropdown choose +Add Custom and search for “Created Timestamp”.You can also find the time stamp from PC CLI using nuclei.For details refer to KB-9045 How to find a VM creation date and time in Prism Central. Always remember to configure adblocking plugin (if you are using one) in the Web browser to exclude Prism from checks. KB-8762 Prism home page shows "loading" when Web browser has AdBlock extension installed
Performance issues usually begin with the scene of an administrator looking at statistics, scratching the back of the head and saying ‘Wait, what?’ The issue can manifest itself in so many ways and the main struggle is to find the entities of the environment that experience issues at the same time and look at the environment layer by layer. Whether the bottleneck lies on the network, compute or storage layer, troubleshooting performance issue is always a tedious and time-consuming task. That is why Nutanix published an article (or two) to help you with the process. Where to look, what to look for and how to interpret the numbers. KB-2345 [Performance] Troubleshooting high CPU in Nutanix environments KB-5012 [Performance] Interpreting vCenter Performance CPU Ready chart values Tech TopX: ESXI Performance Troubleshooting
HI, recently someone upgraded a VM in our Nutanix cluster without documenting it. I know when it happened from the Tasks view in Prism. I need to konw if there is a log of who has logged into Prism to correlate but can’t seem to find this. Any ideas? Thanks!
Prism Pro is a set of additional features that are unlocked when Prism Central is licensed. These features include capacity planning, custom dashboards, and advanced search capabilities. The one-click planning feature that lets you forecast future workload growth so you can expand accordingly to meet the demands.Capacity PlanningThis feature ensures that environment never runs out of capacity. You can forecast shortages of CPU, memory, and storage, defining how much time until a given resource is exhausted. Information is based on machine-learned consumption behavior based on current and past utilization. This also provides a recommendation of scale out options.Just-in-time ForecastingAllows an environment to scale without over or under provisioning. This allows you to define additional workloads to add to an environment and identify whether the environment has resources to handle such workloads. If sufficient resources do not exist, Just-in-time Forecasting can recommend resources to
I”m trying to find out if Nutanix has a scheduled patching regimen similar to Microsoft’s Patch Tuesday. How are admins informed when there is an emergency patch? Do these include patches to Prism, to AOS and firmware patches? Thank you.
Hi,I have 1 cluster 4 nodes nutanix. Any baseline to create a Preventive Maintenance Checklists Document?I just want to create a documentation saying that my nutanix cluster is just fine aside from NCC results. Maybe checklists for physical condition, firmware and etc? Thank you.
In this ever-changing digital world, there is a need to secure and protect digital data. Nutanix provides continuous fixes and updates to address security threats and vulnerabilities (CVEs).To find information about available security fixes and updates, navigate to the “Security Advisories” section within the Support Portal.Nutanix is also committed to providing security patches/releases in a timely manner on any new CVEs that are discovered. The following KB outlines the Nutanix CVE Patching Schedule/Policy: KB 4110
Network Time Protocol (NTP) is a protocol for clock synchronisation between computers. The hosts and CVMs in a Nutanix cluster must be configured to synchronise their system clocks with a list of stable NTP servers. Generally, at least 1 (one), but preferably 3 (three) or more reliable off-cluster NTP servers are configured on the cluster. To avoid split brain scenarios, it is always recommended to configure an odd-number of NTP servers. Some of the primary things when it comes to NTP configuration and troubleshooting can be covered using these quick links. Please note that most of the links here are for the latest LTS version - AOS 5.15(as of July 28th, 2020). To configure NTP servers on the CVMs and AHV, see Configuring NTP Servers in the Prism Web Console Guide. (Configuring NTP servers via Prism will update both the CVMs and the AHV hosts). To configure NTP servers on ESXi hosts, see Configuring Network Time Protocol (NTP) on ESX/ESXi hosts using the vSphere Client (2012069).
Hi everyone Is it possible for future versions of AHV to see the operating system fields of virtual machines?
We are trying to pull VM Creation Date on one of our AHV Cluster, we also have PC in our environment. Kindly advise if any way to fetch the VM Creation date.
Hi Guys I have a multiple of VMs protected, I would like to restore them to a smb. Severe are in different protection domains. Is there a script I can use to restore multiple vms in different PDs at once?
Each of the services in Nutanix Cluster has a definitive purpose. Nutanix ClusterHealth service, for example, is practically an in-build monitoring system of the cluster. Nutanix Cluster Checks (NCC) and pre-upgrade checks rely on ClusterHealth. As with troubleshooting of any service status, the first thing to do is to attempt to start the service manually from the CMV. If that does not help and you are not running Hyper-V on the host check the logs located at /home/nutanix/data/logs/ for any FATAL files and inspect records in there. Note: Foundation process does not run when the Node is a part of a Nutanix cluster. For logs location, handy commands and more tips be sure to check KB-1518 NCC Health Check: cluster_services_down_check. While the article is on an NCC check the troubleshooting steps and commands apply to the ClusterHealth service as much as to other services. KB-6375 Pre-Upgrade Check: test_cluster_status
As your environment grows, keeping track of various entities (VMs, Clusters, services, etc.) can become cumbersome. In order to keep tabs on it, you can utilize a feature called “Entity Exploring” found in Prism Central. With this feature, you can look at data across many clusters which are registered to Prism Central. For example, under the VM->Entities menu, you can create a custom view that lists additional columns (vCPUs, NICs, etc.) to your liking, and then save this view for future use. You can even export this view to a CSV file to save locally. Once you have created this granular view, you can then apply filters to drill-down even further if desired. More examples of the customization of the “Entities” menu can be found in the “Entity Exploring” document of the Support Portal.
Did you know you can restore a deleted VM from a leftover snapshot? If a virtual machine (VM) was accidentally deleted or migrated to a remote site provided the snapshot remains the VM can be recovered. When deleting a VM using Prism an option ‘Delete all snapshots’ is available. When the option is not selected the snapshots remain in the container and become orphaned. Similarly, when a VM is migrated to a remote site an is disassociated with the protection domain the snapshot remains. The recovery operation relies on a clone operation in aCLI. Clone a VM from a snapshot is essentially what happens when recovering the VM. Follow KB-6194 Cloning from AHV Orphaned Snapshots for a simple set of instructions on the process. Prism Web Guide contains more information on Disaster Recovery operations.
I set the security policy named LAMP as shown in the figure, and set it to monitor mode. LAMP-DB CentOS 7 MySQL is running on port 3306. LAMP-WEB database request is accepted. There is no control by fiwewalld. Can VMs in this tier talk to each other ? - No LAMP-WEB CentOS 7 Wordpress and apache are running. Web services are launched on http port. There is no database, no control by fiwewalld. Can VMs in this tier talk to each other ? - No Then I did the following: View LAMP-WEB in browser from 192.168.0.0/23 segment. Ping to LAMP-DB from 192.168.0.0/23 segment. However, “Monitoring” screen shows “Tcp Port:80 No flows found” (as shown in the figure) and despite success of the ping, “No uncaptured traffic flows were detected." is displayed. Why can't "Monitoring" catch the packets?
As a starting point of all or nearly all troubleshooting Nutanix recommends running Nutanix Cluster Checks (NCC). NCC identifies any risks or misses in the cluster configuration, hardware state, software components, networking and more. Checking for presence of orphan VM snapshots is part of NCC which, I think is quite convenient as chasing orphaned snapshots can be a hassle. Should you already know that there is a snapshot that is orphaned it can be removed with a single aCLI command. To be on the safe side always validate that the VM no longer exist. Otherwise, solution section of KB-3752 NCC Health Check: orphan_vm_snapshot_check lists all the steps necessary. More on the subject: NCC - why it is important to upgrade it
Who can say they have never experienced a VM boot issue? We have all been there. Something quick and easy like creating a VM from scratch or powering up one after migration. Annoying but can be solved. On AHV the root cause may be hidden in one of the below: Incorrect boot device selected. Incorrect boot mode is selected: BIOS or UEFI. If the virtual machine was initially created in UEFI boot mode and then boot mode was changed to BIOS or vice versa, then VM will fail to boot. VirtIO drivers are not installed in the guest OS - results in PSOD also known as blue screen of death. VirtIO drivers used to be installed but are removed after a sysprep. You may find verification steps and resolution in the articles below:KB-9400 AHV | Troubleshooting Virtual Machine boot failuresKB-5436 AHV | VirtIO drivers may be removed after OS generalization (sysprep)KB-5338 AHV | How to inject storage VirtIO driver if it was not installed before migration to AHV More on the subject: KB-5666 AHV |
Is there a way to make a UEFI based vm via comandlets or the API? This posting seems to leave those two out. https://next.nutanix.com/prism-infrastructure-management-26/uefi-on-vms-faq-37339
Hi Folks, We have been getting alerts on memory usage from the Prism Central VM. It seems to be a known issue. So to suppress the alerts from repeatedly generating cases in our internal ticketing system I am working on setting up an automated restart when memory usage gets to 90%. I’ve written the code to do this (a mix of the Nutanix cmdlets and Rest API calls in PowerShell). It works fine when run under my own account. My account has the ‘User Admin’ role. I should also mention we map roles to AD groups. However I wish to move the scheduled job to run under the context of a service account. The service account has been assigned a Custom Role with a single assignment ‘Update VM Power State’ I’m added the Role to my Service account user and applied it to the Prism VM (and a Test VM). I’m wondering if I need to grant more permissions as I am getting an error Access is deniedSet-NTNXVMPowerState : The remote server returned an error: (403) Forbidden.
Microsoft has enabled LDAP channel binding and LDAP signing in March this year in Active Directory Windows Servers architecture. Nutanix recommends changing Prism Authentication from LDAP on port 389 to LDAPS on ports 636 or an SSL encrypted port 3269. Two things to note with the process are: changing just the port number is not enough because the LDAP protocol also needs to change to LDAPS Prism self-signed certificates work with the LDAPS so no extra hassle The process is straightforward and only requires a change of a URL syntax in Prism settings. For instructions and verification steps see:KB-9029 Changing LDAP port 389 authentication to Secure LDAP (LDAPS) ports 636 or 3269. More on LDAP:KB-3363 Prism: Troubleshooting LDAP and AD Issues for Prism Log On For more Information about Microsoft change:2020 LDAP channel binding and LDAP signing requirements for Windows.
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.