Get guidance, share wins, and ensure smooth Nutanix deployments.
Recently active
In some cases, you might have to permanently remove a physical node / host from a Nutanix cluster. There are two scenarios in node removal. Permanently Removing an online node Removing an offline / not-responsive node in a 4-node cluster, at least 30% free space must be available to avoid filling any disk beyond 95%. You cannot remove nodes from a 3-node cluster because a minimum of three Zeus nodes are required. Some Points to consider before initiating node removal: Sufficient Disk space available on other nodes in the cluster User Virtual Machine relocation (if required) Any software upgrade should not be running Checklist on verifying cluster health status Data resiliency is “OK” (green) in Prism Run a complete “ncc report” either from prism or CVM cli: ncc health_checks run_all Depending on the size of data, node removal can be lengthy process, which involves relocating data from the node to other healthy nodes in the cluster. Node removal also remo
Hi,I am implementing a Nutanix Cluster with Lenovo HX5520, when I register the Prism in vCenter the process appears as completed but I cannot create virtual machines by Prism Element and when running the NCC I return communication errors with the vCenter is not stabilized. I did the communication test of the CVMs with the vCenter on ports 80 and 443 and the connection worked successfully. when checking by cli I see that the connection is ok, but it does not provide me with the settings of vcenter follows the difference from another implementation any idea what it might be?
Below are new knowledge base articles published on the week of December 27, 2020-January 2, 2021.KB 10019 - Alert - A110024 - AwsDefaultUVMSecurityGroupNotFound KB 10507 - Nutanix Move | VM migration fails if Hyper-V VM has Fibre Channel disk controller attached KB 10528 - Era - Operation failed with Internal Error after changing Cluster Account Password KB 10529 - [ Karbon ] How to configure email alerts for a Karbon kubernetes cluster KB 10530 - NCC-4.0.0 : Health Server logs might fail to rotate and fill up /home partitionNote: You may need to log in to the Support Portal to view some of these articles.
Below are the top knowledge base articles for the month of December 2020.KB 7503 - NX Hardware [Memory] – G6, G7 platforms - DIMM Error handling and replacement policy KB 10475 - LCM 2.4 inventory failure - [SSL: UNKNOWN_PROTOCOL] unknown protocol (_ssl.c:618). Not fetching available versions for module KB 4141 - Alert - A1046 - PowerSupplyDown KB 1540 - What to do when /home partition or /home/nutanix directory on a Controller VM is full KB 1113 - HDD/SSD Troubleshooting KB 4409 - LCM: (LifeCycle Manager) Troubleshooting Guide KB 4158 - Alert - A1104 - PhysicalDiskBad KB 2090 - AHV host networking KB 4519 - NCC Health Check: check_ntp KB 2473 - NCC Health Check: cvm_memory_usage_check KB 4116 - NX Hardware [Memory] – Alert - A1187, A1188 - ECCErrorsLast1Day, ECCErrorsLast10Days KB 4273 - NCC Health Check: aged_third_party_backup_snapshot_check and aged_entity_centric_third_party_backup_snapshot_check KB 6945 - How Upgrades Work at Nutanix KB 1863 - NCC Health Check: sufficient_disk_s
What is happening to my 2 node cluster during a failover or an upgrade?What does a recovery process look like after a node failure?If you are wondering the above, we have the answer for you! You can monitor the progress of your 2 node cluster in these situations through Prism Element.To monitor node recovery progress after failover: Registering a witness is highly recommended to help the cluster handle the failover situation automatically and gracefully. Stand-Alone mode: A failed node would trigger cluster to transition into stand-alone mode during which the following occurs: Failed node is detached from metadata ring. Auto rebuild is in progress. Surviving node continues to serve the data. Heartbeat: Surviving node continuously pings its peer. As soon as it gets a successful reply from its peer, clock starts to ensure that the pings are continuous for the next 15 minutes. If a ping fails after a successful ping, the timer will be reset. Prism Element Home page shows Critical
Prism Central includes machine-learning capabilities that analyze resource usage over time and provide tools to monitor resource consumption, identify abnormal behavior, and guide resource planning. These tools include VM "right sizing" where VMs are analyzed and those that exhibit inefficient profiles are identified. Anomaly detection to record when performance or resource usage is outside an expected range based on learned VM baseline behavior. "Smart" alerts that trigger when specified anomalies are recorded. Reports that summarize cluster efficiency. VM Right SizingIt is useful to look at the profile of your VMs when analyzing problems in a cluster or assessing future resource needs. This can help you identify VMs that are not optimally configured such as ones that consume too many resources, are constrained, are over provisioned, or are inactive.Anomaly Detection:The right sizing feature identifies inefficient VMs that fit one of the profiles described as below: Bully VM : A
Nutanix takes a holistic approach to security with a secure platform, extensive automation, and a robust partner ecosystem. The Nutanix security development life cycle (SecDL) integrates security into every step of product development, rather than applying it as an afterthought. The SecDL is a foundational part of product design. The strong pervasive culture and processes built around security harden the Enterprise Cloud Platform and eliminate zero-day vulnerabilities. Efficient one-click operations and self-healing security models easily enable automation to maintain security in an always-on hyperconverged solution.Since traditional manual configuration and checks cannot keep up with the ever-growing list of security requirements, Nutanix conforms to RHEL 7 Security Technical Implementation Guides (STIGs) that use machine-readable code to automate compliance against rigorous common standards. With Nutanix Security Configuration Management Automation (SCMA), you can quickly and continu
Below are new knowledge base articles published on the week of December 20-26, 2020.KB 10328 - Windows VM on AHV with Nutanix VirtIO Unable to Read "Physical Disk Serial Number" Intermittently KB 10349 - DHCP-client startup impacting Windows VM guest services KB 10415 - NX-8170-G7 imaging fails with Foundation < 4.5.4 KB 10446 - Cannot provision node due to AWS Quota exceeded issue. Quota type cpu. KB 10456 - ESXi 6.5 failure when imaging with Foundation 4.5.4.2 KB 10474 - Objects - Manual steps to configure emails for Objects Alerts KB 10487 - Unable to create thick provision disks from Nuranix NFS datastore on VMware due to nfs-vaai plugin missing on ESXi hosts KB 10508 - Rack aware settings not working when PE launched from PCNote: You may need to log in to the Support Portal to view some of these articles.
Here I discuss the effects of Enabling or Disabling Deduplication on a container even if the Container has data already written to it. The benefit of Compression and Fingerprinting+Deduplication is to hold more data in the container, by reducing the stored size and avoiding duplicate data, respectively.Nutanix’s intelligent selection of dedupable candidates prevents deduplication being performed where the benefit would be low. Deduplication Best Practices: Enable deduplication Do not enable deduplication Full clones Physical-to-virtual (P2V) migration Persistent desktops Linked clones or Nutanix VAAI clones: Duplicate data is managed efficiently by DSF so deduplication has no additional benefit Server workloads: Redundant data is minimal so may not see significant benefit from deduplication Enabling Dedupe:Fingerprinting is method of creating signatures of the data in Metadata. Fingerprint-on-write (Cache-Ti
I am moving virtual machines with Move 3.6.2 from Hyper-V to an AHV CE cluster. Everything goes well, but after the cutover the AHV VM is stuck at the ‘Press F2 for EFI boot manager’ screen and cpu usage around 52%. Pressing FN+F2, F2, CTRL+F2, ALT+F2 or other combinations have no effect. Please advise. Hyper-V host is Win2019, VM is 2019 with boot from bootmgfw.efi and secure boot disabled. Also tried with Edge, Chrome and IE, same result, stuck at ‘Press F2 for EFI boot manager’ screen.
Have you guys been utilizing the Analysis charts feature effectively? It gives you the ability to create charts that can monitor a variety of performance metrics over a week, month, or a custom time range. Here are a few ideas on how we resolved some issues seen in the field. Memory Usage (%) - Create a chart to track memory consumption or one or multiple VM’s over a time interval Hypervisor CPU Usage (%) - Measure the CPU usage of one or more hosts over time and gain a better understanding of resource constraints if any Storage Controller Latency - This is particularly useful if you suspect performance issues. Creating charts ranging to a few weeks back gives you a comparative analysis and a benchmark for the current latency observed Storage Container Usage - A classic use case for this is when multiple VM’s have been migrated out of the cluster over many days but you suspect the storage space has not been reclaimed Replication Bandwidth - Transmitted - If there are
This reference covers the v1 Nutanix API. The complete reference for the v2 Nutanix API, including code samples in multiple languages, and tutorials are available at http://developer.nutanix.com/ Users Get Logged In Users DetailsGET /users/logged_in_users Get Logged In Details of a userGET /users/logged_in_users/{userName} Get Logged In Users DetailsGET /users/logged_in_users Get Logged In Details of a userGET /users/logged_in_users/{userName} Get Logged In Users DetailsGET /users/logged_in_users path /users/logged_in_users method GET nickname getAllLoggedInUsersInfo type get.base.EntityCollection<get.dto.auth.UserDTO> Property Type Format entities array errorInfo get.base.ErrorInfo metadata get.base.Metadata Get Logged In Details of a userGET /users/logged_in_users/{userName} path /users/logged_in_users/{userName} metho
Below are new knowledge base articles published on the week of December 13-19, 2020.KB 8562 - NCC Health Check: robo_witness_configured_check KB 8563 - NCC Health Check: robo_witness_state_check KB 8565 - NCC Health Check: robo_cluster_witness_sync_check KB 9271 - NCC Health Check: ahv_fs_integrity_check KB 9472 - NCC Health Check: category_protected_vms_multiple_fault_domain_check KB 9525 - Alert - A200330 - Prism Central home partition expansion check KB 9713 - Alert - A130340 - MetroConnectivityUnstable KB 9716 - NCC Health Check: stale_synchronous_replication_parameters_check KB 9845 - NCC Health Check :- "file_server_cvm_config_check" KB 9988 - Pre-Upgrade Check: test_if_expand_cluster_is_not_in_progress KB 10000 - NCC Health Check: objects_deployed_on_unsupported_pe KB 10248 - Alert - A130340 - Cross-container disk migration task is paused. KB 10323 - Move VMs from Protection Domain to Category for Leap KB 10339 - Skipping application consistent snapshot for VM with NVMe disks KB
Hello i have a 4 Lenovo HX 3320 nodes. I updated all bmc on this servers. When i am installing new aos throw Foundation 4.6 its fails with this error Hele is log 2020-12-10 12:38:39,191Z DEBUG Setting state of <ImagingStepValidation(<NodeConfig(192.168.1.22) @e2d0>) @e5f0> from PENDING to RUNNING2020-12-10 12:38:39,197Z INFO Running <ImagingStepValidation(<NodeConfig(192.168.1.22) @e2d0>) @e5f0>2020-12-10 12:41:01,398Z DEBUG Cache HIT: key(<function common_validations at 0x03D71230>_()_{'global_config': <foundation.config_manager.GlobalConfig object at 0x054D0E90>})2020-12-10 12:41:01,405Z DEBUG Setting state of <ImagingStepValidation(<NodeConfig(192.168.1.22) @e2d0>) @e5f0> from RUNNING to FINISHED2020-12-10 12:41:01,410Z INFO Completed <ImagingStepValidation(<NodeConfig(192.168.1.22) @e2d0>) @e5f0>2020-12-10 12:41:01,415Z DEBUG Setting state of <GetNosVersion(<NodeConfig(192.168.1.22) @e2d0>) @e550> from PENDIN
LCM (Life Cycle Manager) is a tool provided by Nutanix to upgrade firmware and software across a Nutanix environment. Depending upon the type of firmware that needs to be upgraded, a reboot of a node may be required.While reviewing the inventory of software/firmware that needs to be upgraded, you may not know which upgrades require a reboot of a node (or not). This information might be crucial in determining the impact and precautions that need to be taken in planning for the upgrade event (change windows, time allocated, etc.).KB 6107 details which upgrades require a node or CVM reboot and which ones do not. Navigate to the
Many users are unaware that there are additional (beyond what is presented via the Prism user-interface) security parameters that can be employed on AHV hosts to increase the overall security of them. These security parameters are configured via Nutanix Command-Line Interface (NCLI) and include the following: Advanced Intrusion Detection Environment (AIDE) - a file and directory integrity checker High Strength Password Enforcement - configure the maximum and minimum number of characters the password must contain along with number of passwords retained in history to prevent repeated use Core Dumps - the recorded state of the working memory for a process is dumped to a file if the process ever crashes Login Banner - display a customized messages when user login to a node More information regarding these parameters, including the procedures to enable/disable them, can be found within the Hardening AHV section of the Nutanix Security Guide. Also to note, there are similar parameters
SAP helps customers migrate from traditional relational databases to their in-memory SAP HANA database to gain more agility in their business processes. Many SAP customers are searching for ways to deploy SAP HANA in an efficient, simple way that minimizes risk while preserving the benefits of an agile platform. Nutanix provides such an option. The native Nutanix hypervisor, AHV, and Nutanix enterprise cloud OS software are certified for production SAP HANA deployments. HCI for SAP HANA CertificationThe certification has two primary segments. 1. As the first step, a platform vendor (Nutanix, in this case) must validate their platform, which consists of a hypervisor and an HCI component.2. In a second step, the hardware OEM must certify a suggested configuration through some additional HCI-related tests. When both parts of the validation are complete, the solution is certified and listed in the HCI for SAP HANA category on the SAP website. The hardware OEM is then responsible for selli
Any modern environments consist of multiple layers each of which contains multiple components. There are switches and routers, firewalls, physical server, application servers, applications themselves and, of course, users. Each of the components has logs of more than one kind, location and severity. All the components interact with each other directly or indirectly. I am certain you have found yourself in a situation where to establish a root cause you had to inspect logs of more than one entity. Establishing a timeline of events is always easier when the sources of the events’ clocks are synchronised and are located in one central location. While the clocks are handled by the NTP the centralised logs location is a syslog server or in this case a remote syslog server implying that it is separate to the origination of the logs. In addition to the benefits already mentioned, remote syslog server allows to access logs for the systems that are already dead, decommissioned or replaced. Nuta
These days I’m managing, for one of customer, the migration of its virtual infrastructure to a Nutanix AHV farm from an old ESXi. It has some old VMs, with Windows XP as guest OS, that have to be migrated in Nutanix. Those VMs have installed custom software that cannot be installed in newer versions of Windows because they have been written in old VB6, so my question are:Can they be migrated over AHV? If yes, can I manage the migration with Nutanix Move or there is another way?
Use the Data Services IP method for external host connectivity to VGs. For backward compatibility, you can upgrade existing environments non disruptively and continue to use MPIO for load balancing and path resiliency. For security, use at least one-way CHAP. Leave ADS enabled. (Enabled is the default setting.) Use multiple disks rather than a single large disk for an application. Consider using a minimum of one disk per Nutanix node to distribute the workload across all nodes in a cluster. Multiple disks per Nutanix node may also improve an application’s performance. For performance-intensive environments, we recommend using between four and eight disks per CVM for a given workload. Use dedicated network interfaces for iSCSI traffic in your hosts. Place hosts that use Nutanix Volumes on the same subnet as the iSCSI data services IP. Use a single subnet (broadcast domain) for iSCSI traffic. Avoid routing between the client initiators and CVM targets. Receive-side sca
Hi, I have 2 Blocks (2 Nodes each) Cluster. I just added 2 DIMM Memory (2 x 32GB) in 1 of my Node. I put 2 memory on slot P1-DIMMC1 and P1-DIMMF1Before Adding DIMMAfterr Adding DIMMAfter that i ran NCC and shows warning like thisAlert from Nutanix ClusterAm i did something wrong with the steps by steps? Since in PRISM, total memory has been increased and i thought it’s just fine.
Below are new knowledge base articles published on the week of December 6-12, 2020.KB 10285 - Nutanix Files :- Alert "Error in updating CA Chain to file servers: Subtask failed for <FSVM Name>" KB 10322 - Alert - A200602 - MicrosegmentationControlPlaneFailed KB 10380 - Xi Leap: VPN troubleshooting guide KB 10400 - Mismatch in Prism Central End of Life alerts KB 10411 - [ Karbon ] K8s clusters using XFS storageclass might encounter mount issues before CSI 2.2 KB 10424 - Performing HYCU Restore workflow for VM protected by Protection Domain or Protection Policy causes DR service crash issues KB 10427 - Identifying CVE and CESA patches in Nutanix ProductsNote: You may need to log in to the Support Portal to view some of these articles.
In many ways, identifying the problem is harder than solving it. At least in IT.In order to better understand the performance of Nodes in a cluster and User VMs, ESXTOP for ESXi hypervisors and TOP for Linux OS based machines provide us an immediate birds eye view of the performance of the host. It is important to understand the output as a tool towards problem-solving. Here I aim to discuss scenarios you might need to isolate and troubleshoot High CPU observed at the Hypervisor or User VMs.ESXTOP : This command when run lists live CPU statistics specific to the Node the command is run on.esxtop outputIf you press “M” it shows you memory metric and “N” for network etc. As Always ‘H’ is for help. We will focus on a few CPU statistics:The output will show you all the VMs and the following metrics corresponding to them.Some Important ones discussed below:%USED, %RDY, %CSTP , %MLMTD and %SWPWT Note:To convert CPU ready % value to ms(milliseconds)CPU ready % = ((CPU summation ready value i
PCI device enumeration is one of a large number of concepts that transitioned from the physical world. Originally, being a bus and a slot number now, of course, is a virtualised concept.When OS boots during the POST process the local devices are enumerated (checked for presence and size). If an expected device is found at the PCI slot that is a failure of the test. In AHV deployments, the Controller VM (CVM) runs as a VM and disks are presented using PCI passthrough. That means that if AHV PCI devices enumeration has changed CVM may not be aware of the fact and attempt to direct I/O to the devices that are no longer accessible via known PCI bus and slot. This can happen after hardware replacement on an AHV host and the CVM will not boot or will boot with only part of the expected devices accessible. Compare PCI slot numbers of SCSI controllers on AHV host and the CVM and update the enumeration of the devices on the CVM. For more details look at KB-7154 AHV | CVM might not boot after ha
Anyone have an issue where you are not able to apply labels anymore to VMs after upgrading to 5.18.1.1 AOS and Prism Central pc-2020.9.0.1? I am unable to apply labels after upgrade.
Already have an account? Login
No account yet? Create an account
Enter your E-mail address. We'll send you an e-mail with instructions to reset your password.