Have questions about how the Nutanix Platform works? Looking to get started - start here!
Recently active
SAS is the leader in analytics and Nutanix is the leader in invisible infrastructure. Nutanix has thousands of customers and many of them already have SAS software running in their organization. They have experienced the benefits of invisible infrastructure and are moving more of their applications (including SAS) to Nutanix. It’s easy to deploy and manage SAS 9.4 and SAS Viya on Nutanix Today, SAS 9.4 helps discover insights, manage data and make analytics approachable. SAS 9.4 has been tested on Nutanix NX Models – both as a hyper-converged infrastructure (HCI), as well as using Nutanix as back-end storage only, for external hosts. Nutanix AHV clusters perform well in both scenarios. SAS software makes great demands of IT infrastructure, so you must get the design right to ensure a successful deployment. Evaluating the SAS I/O requirements accurately is pivotal. Nutanix encourages involvement of Nutanix engineers to determine the back-end service requirements as well as optimal imple
For most maintenance tasks and upgrades we can keep the cluster up and VMs running, but in some cases the whole cluster will need to be shut down. If you just need to power off a single node, a cluster of three or more nodes won’t need to stop. To stop a single-node cluster please see the section “Shutting Down a Single-node Cluster” in the NX and SX series hardware administration guide. To stop a single node in a larger cluster see the section “Shutting Down a Node in a Cluster (AHV)” If there's going to be a site power outage, a full network outage, or physical relocation of the whole cluster you're going to want to gracefully shut down the whole cluster. The full procedure is covered in the article Shutting Down an AHV Cluster for Maintenance or Relocation. In summary the procedure will be as follows: Update NCC and perform a health check, then address any items of concern. Shut down all the user VMs. Stop any Nutanix Files cluster, if applicable. At this point no VMs other than
Let’s say that you ran the health checks on your cluster and received a failure under the component “cvm_name_check”, what does it mean and how do you fix it? The NCC health check cvm_name_check ensures that any renamed CVMs (Controller VMs) conform to the correct naming convention to avoid issues with certain operations that depend on identifying the CVM from UVMs on the same host. The default Controller VM naming is NTNX-<block_serial>-<position-in-block>-CVM. The display name of the Controller VM must always: Start with "NTNX-"; and End with "-CVM" For more information check out: https://portal.nutanix.com/#/page/kbs/details?targetId=kA00e000000XfCMCA0 To see how to modify the hostname of the Controller VM check out: https://portal.nutanix.com/#/page/kbs/details?targetId=kA032000000TUjkCAG You followed the naming convention and the check is still showing a failure? This might be a false-positive alert depending on your AOS, contact Nutanix support for verificatio
You may have noticed when adding disks to the node or replacing disks with larger capacity ones, utilisation distribution between the disks does not occur immediately. You check on the cluster sometime later and notice that newly added disks still show minimal usage, much lower than expected. By default, the aim is to bring disks utilisation within +/-7.5% spread of the tier utilisation. There are some things to consider when expecting a certain outcome: Disk balancing is not triggered unless the tier usage is at least 35%. Only 1 GB of data is moved during a Curator scan per node. Even if the tier usage is below 35% should any disk usage across the cluster reach 70%, disk balancing takes place. Disk Balancing: Disk balancing ensures data is evenly distributed across all disks in a cluster. In disk balancing, data is moved within the same tier to balance out the disk utilization. This is different from ILM (Information Lifecycle Management), where data is moved between dif
Would like to put together a report that would in detail break down the storage used by VM snapshots, as well as, storage used by Protection Domains. Have been digging around and found a couple scripts that display some useful information however not quite what I’m looking for. Anyone have an idea on how best to proceed?
When a Nutanix / vSphere cluster is deployed by Foundation the recommended drivers are installed, but after some time you may want to check if there is a newer driver recommended. From the Nutanix perspective, we have covered this with an NCC Health Check: esx_driver_compatibility_check so if you update NCC and run a health check, this check should tell you whether there is a later driver version qualified by Nutanix. To run the check from the CLI use “ncc health_checks hypervisor_checks esx_driver_compatibility_check” from any CVM in the cluster. You may see a newer driver listed for your NIC hardware and ESXi version. A newer driver may not have been qualified yet by Nutanix and in some cases could cause issues for the cluster, so generally we recommend staying with the recommended drivers as identified by NCC.
Q. Does Nutanix support inline encryption? Inline encryption is not currently supported on the Nutanix platform. However Data At Rest Encryption (DARE) of two kinds is supported: Using Self Encrypted Drives (SED) is supported Security Guide v5.16: Preparing for Data-at-Rest Encryption (SEDs) Security Guide v5.16: Configuring Data-at-Rest Encryption (SEDs) Software Only Data Encryption Security Guide: Data-at-Rest-Encryption (Software Only) Q. How do I know that the data is encrypted? More details around the encryption status and logs can be viewed via the nCLI, using REST APIs or PowerShell cmdlets. KB-7846 How to verify that data is encrypted with Nutanix data-at-rest encryption Q. Is it possible to monitor the encryption? Monitoring of the encryption state is done via our Nutanix Cluster Checks (NCC) that generate an alert on any issue detected within the cluster. Please keep in mind that enabling encryption is a cluster-scope setting. Q. What is recommended sizing of the CV
Below are new knowledge base articles published on the week of March 15-21, 2020. KB 8885 - Alert - A15039 - IPMI SEL UECC Check KB 9009 - AHV | No Intel Turbo Boost frequencies shown in the output of "cpupower frequency-info" command KB 9070 - [CSI] PVC Volumes Stuck in Pending State | Error: Secret value is not encoded using '<prism-ip>:<prism-port>:<user>:<password>' format KB 9071 - Configuring hypervisor after satadom replacement fails with phoenix 4.5.2 KB 9085 - [Objects 2.0] Error creating the object store at the deploy step. KB 9095 - Nutanix Files- FSVM expansion may file with error IP already in use KB 9102 - How to identify plugged-in SFP module hardware details KB 9103 - WARN: Could not use proxy. URL Error <urlopen error [SSL: TLSV1_ALERT_INTERNAL_ERROR] Note: You may need to log in to the Support Portal to view some of these articles.
In a Nutanix AHV cluster the image service is used to index and manage ISO and virtual disk images for cloning to new VM disks or mounting to the virtual CDROM. With the addition of Prism Central 5.5 or later, this adds a global image service to manage these files across multiple clusters. When managing images from Prism Central we sometimes will see an image show up on a cluster as "inactive". This means the metadata for the image exists but the file does not exist locally on that cluster. The article "Prism Central: Adding Images to Prism Central" gives a few options to remediate this condition when an image is needed on a certain cluster but is inactive. These methods are useful with Prism Central 5.5 and 5.10 versions. In Prism Central 5.11 we have added image placement methods to control where your images will be available. During image upload you can choose to select individual clusters where the image should reside, or you can apply a category to the image to utilize an imag
What is Erasure Coding? Erasure coding increases the usable capacity on a cluster. Instead of replicating data, erasure coding uses a parity information to rebuild data in the event of a disk failure. The capacity savings of erasure coding is in addition to deduplication and compression savings. If you have configured redundancy factor 2, two data copies are maintained. For example, consider a 6-node cluster with 4 data blocks (a b c d). In this example, we start with 4 data blocks (a b c d) configured with redundancy factor 2. In the following image, the white text represents the data blocks and the green text represents the copies. Data copies before Erasure Coding Computing Parity Data copies after Computation of Parity Erasure Coding Best Practices and Requirements: A cluster must have at least four nodes populated with each storage tier (SSD/HDD) represented to enable erasure coding. Avoid strips greater than (4, 1) because capacity savings provide diminishing returns and
What is Metro Availability? Nutanix provides native “stretch clustering” capabilities which allow for a compute and storage cluster to span multiple physical sites. In these deployments, the compute cluster spans two locations and has access to a shared pool of storage. The solution is currently available for ESXi only. This expands the VM HA domain from a single site to between two sites providing a near 0 RTO and a RPO of 0. In this deployment, each site has its own Nutanix cluster, however the containers are “stretched” by synchronously replicating to the remote site before acknowledging writes. The following figure shows high-level architecture of a Nutanix Metro Availability deployment: The following figure shows an example link failure: Nutanix Metro Availability also can be set up with Async-DR replication to a third site to combine the multi-site resiliency of the Metro setup with traditional space-efficient incremental snapshot backups. Metro Availability Configuration
Is there a way to execute a script (or PowerShell script/command) inside a newly created AHV VM which does not have an IP address yet? The VM has Nutanix Guest Tools installed. I’m looking for automating the creation of a new VM. In VMware, I use the Invoke-VMScript command. I didn’t know if Nutanix had something similar? Thanks! -TimG
This alert is generated when an API comes in which is authenticated as "admin". Nutanix recommends any script or third party application sending APIs to the cluster should use a service account rather than using 'admin'. You can read more about this alert from the article "Alert - ExternalClientAccessCheck" If you are seeing this alert, it is informing you that some system is authenticating as admin. To aid in investigation the IP address is provided. The intent is that any 3rd party application or script should be using a service account and not ‘admin’ as this makes command auditing much more reasonable and helps keep the admin password secure. When a third party application such as Veeam is set up to authenticate to the cluster as ‘admin’ that should generate this alert. If you log in to Prism Element as admin, access the REST API explorer, and then test an API you should see this alert because that’s your desktop sending an API as ‘admin’. Likewise if you set up a PowerShell script
All, I have a 6 Node 1065 system in 2 blocks. Recently one of the CVMs (node 2, block A) crashed. When rebooted it was just going in loops. When diagnosed it seems the SSD (not the SATADOM) had failed and we replaced it. When we try to boot the CVM, it still just loops. We were told to boot that node with Phoenix which the cluster provided me for download. I do that and it doesn’t load Phoenix and gets errors instead. I’m looking for a suggestion of how to get the node back to 100%. At this point (and throughout) the ESXi on the SATADOM has booted fine and I guess if I didn’t care about the storage side I could just ignore this but I’d like the system to be fully healthy. Any suggestion about how to get the CVM working again would be appreciated. Thank you Johan
Below are new knowledge base articles published on the week of March 8-14, 2020. KB 7424 - NCC Health Check: metro_invalid_break_replication_timeout_check KB 9000 - NCC reports LSI firmware is blacklisted for DELL XC nodes, when LCM inventory does not have any newer versions KB 9042 - Prism Central session times out unexpectedly when a user logged in with 'Admin' role. KB 9063 - [Karbon] Kube DaemonSet Rollout Stuck Alert; Daemonset wrongly reports unavailable pods KB 9068 - AHV | nutanix-network-crashcart scripts fail with "No module named fc_progress" error on hosts imaged with Foundation 4.5.2 KB 9074 - [Karbon] Kubernetes Upgrade Fails with Error: Upgrade failed in component Monitoring Stack Could not upgrade k8s and/or addons Note: You may need to log in to the Support Portal to view some of these articles.
Below are new knowledge base articles published on the week of March 1-7, 2020. KB 8869 - NGT installation via Prism Central on Windows Server 2016 or more recent Operating Systems fails with INTERNAL ERROR message KB 8905 - How to download images from Prism Element clusters via command line KB 8917 - NCC - ERR : The plugin timed out KB 8993 - Foundation : Upgrade foundation using LCM Dark site bundle KB 8997 - Not able to delete a Role in Prism Central KB 8998 - Era Registration fails if container is not mounted on all hosts KB 9004 - Increased number of connections to File Server once migrated from Windows to Nutanix Files KB 9013 - How to Create a Shared Folder in Windows Server 2016/Windows 10 KB 9016 - Unable to open Java console after BMC upgrade from version 7.00 to 7.05 KB 9028 - ERA-DB provisioned from the OOB template failed to register with ERA server KB 9031 - Prism and Microsoft LDAP Channel Binding and Signing KB 9045 - How to find a VM creation date and time in Prism Cen
2 of 3 nodes are fine and working. The 3nd CVM is up and i could ping it. Restart the Cluster with “allssh genesis stop cluster_health; cluster start” does not start these Cluster Partner. After Login in Prism i saw a “Disk degraded” for these 3rd node and now these disk is missing. How to fix node Nr. 3 in a 3 Node Cluster?
Suppose you need more disk capacity on a virtual machine in your environment. You choose the VM in Prism, click ‘update’, select to edit the appropriate disk, and change the size of the disk from 200 to 300GiB. You click update and see that the task completes successfully, then close the VM update UI. The VM details reflect the increase in disk space, but when you access the VM it appears the capacity of the drive is unchanged! This is actually expected. There is just a bit more work to be done. The partition will need to be extended following the steps for your VM guest operating system. You can see the steps to complete this in Windows from the KB article “Expand volume group disk size on Windows OS” or if you are using Linux, check the KB article “Increase disk size on Linux UVM”
Let’s say you received an alert stating that all CVMs are not in the same timezone or all hosts are not in the same timezone. What does it mean? Well, as simple as the alert indicates, the CVMs/hosts are not in the same timezone. We need to ensure that the same timezone is configured across all the CVMs/Hosts as it ensures that all the guest VMs log messages are timestamped consistently. How will you know about the timezone issue? There is an NCC health check, “same_timezone_check” in place to inform any discrepancy in the timezones. To know more about the alerts and errors which can be seen and how to change the timezone, take a look at https://support-portal.nutanix.com/#/page/kbs/details?targetId=kA0600000008hm9CAA Have any questions? Leave a comment and let’s start a discussion.
First, let us understand what NTP (Network time protocol) is. An NTP server is a time server that is used to keep/sync the time in your cluster. An NTP server can be public or private depending on the strictness of your environment. To know how to configure NTP in your Nutanix cluster, take a look at- https://support-portal.nutanix.com/#/page/docs/details?targetId=Web-Console-Guide-Prism-v5_16:wc-system-ntp-servers-wc-t.html After the NTP server is configured, the genesis leader becomes the NTP leader, which means that the genesis leader is syncing time to the NTP server and other CVMs are syncing time with the genesis leader. How NTP works in AHV:- It’s as simple as it gets. The AHV hypervisor takes the same server configured on the cluster and syncs the time with it individually. There are no extra steps required to configure the NTP server on the AHV hosts. How NTP works in ESXi:- The ESXi cluster does not take the server configured on the Nutanix cluster and it needs to be
I am having an External NTP server, Which is tagged to my CVM, unfortunately my AHV is not corresponding with mt NTP server even it is reachable from my hypervisor level. I have checked my Hypervisor thru ssh and run the command ‘date’ its gives me different time from my NTP server.
You’re probably aware all the CVMs and hypervisor hosts in a cluster need to have IPs in the same subnet, but what about the IPMI? What are the requirements, and what’s involved in changing the IP? The requirements are actually quite flexible. The IPMI does not have to be in the same subnet as the hosts and CVMs. You’ll see an alert and a cluster health warning if you don’t restart the genesis service on the CVMs after making changes, but the configuration can be whatever works best for your organization. To restart the genesis service, log into the CVM via SSH or console as the user ‘nutanix’ and run the command “genesis restart”. This restart of the genesis service is non-disruptive. Actually, you can plug the IPMI into an isolated network or configure it on a different VLAN. You could even leave it unplugged when not in use. The IPMI is very useful for installs and updates, and for troubleshooting hardware issues or an unexpected reboot but it is not required for day to day operati
Is Windows Server Failover Cluster supported with shared vmdk with esx? Or only Nutanix Volumes with iscsi? Hypervisor is esx.
What is the recommended maximum storage utilization in a cluster? Customers can observe cluster issues when they use more than 90 percent of the total available storage on the cluster. Here is a brief explanation on the recommended storage utilization in respect to the replication factors 2 & 3. For a cluster to be considered healthy and functioning as expected, the cluster has to tolerate at least one node failure for data resiliency. Here is an explanation on how much free space you need and how to calculate it. The formula for calculating the maximum recommended usage for clusters is one of the following: Recommended maximum utilization of a cluster with containers using replication factor (RF)=2 M = 0.9 x (T - N1) Recommended maximum utilization of a cluster with containers using replication factor (RF)=3 M = 0.9 x (T - [N1 + N2]) M Recommended maximum usage of the cluster T Total available physical storage capacity in the cluster N1 Storage capacity of the node with t
Hi, I want to get these numbers with API call. What is the best way to achieve this? Thanks
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.