Loading...
Loading...
Clusters with PowerProtect Data Manager (PPDM) workflow that were upgraded to OneFS 9.10.1.8 may experience NFS data unavailability and NFS server not responding errors. Restarting NFS brings no relief, as it is the NFS restart process (specifically, the NFS exports refresh portion of the restart) which causes the issue in the first place. The /var/log/nfs.log file on the PowerScale nodes will be observed to be constantly rolling over with errors similar to this example: 2026-07-16T08:41:27.011083-04:00 <30.4> PowerScale-1(id1) nfs[34006]: [nfs] Retryable error inserting "/ifs/PowerScale/.snapshot/DELL-1727324919512335885" for export 3127: 0xc0000034(STATUS_OBJECT_NAME_NOT_FOUND), will schedule zone refresh 2026-07-16T08:41:27.011098-04:00 <30.3> PowerScale-1(id1) nfs[34006]: [nfs] Refresh Required for zone 7, export: 3127, status: 0xc0000467 (STATUS_FILE_NOT_AVAILABLE) 2026-07-16T08:41:27.057679-04:00 <30.3> PowerScale-1(id1) nfs[34006]: [nfs] Failed to insert "/ifs/PowerScale/.snapshot/DELL-1727349964027407987" for export 3155 with status 0xc0000034(STATUS_OBJECT_NAME_NOT_FOUND) 2026-07-16T08:41:27.057701-04:00 <30.4> PowerScale-1(id1) nfs[34006]: [nfs] Retryable error inserting "/ifs/PowerScale/.snapshot/DELL-1727349964027407987" for export 3155: 0xc0000034(STATUS_OBJECT_NAME_NOT_FOUND), will schedule zone refresh 2026-07-16T08:41:27.057715-04:00 <30.3> PowerScale-1(id1) nfs[34006]: [nfs] Refresh Required for zone 7, export: 3155, status: 0xc0000467 (STATUS_FILE_NOT_AVAILABLE) 2026-07-16T08:41:27.101348-04:00 <30.3> PowerScale-1(id1) nfs[34006]: [nfs] Failed to insert "/ifs/PowerScale//.snapshot/DELL-1779496238538318421-d240e353-e848-5dcb-954f-463c0d998ca0" for export 84297 with status Isilon NAS Backups on PPDM may fail continuously with the following error due to PPDM not getting a response from the NFS service on PowerScale: "Could not contact NFS server to list exports by path: Encountered error status STATUS_NOT_FOUND (0xc0000225). Network File System (NFS) protocol becomes unresponsive to specific client requests, which may result in data unavailability. NFS clients may experience "NFS server not responding" errors when accessing data on the PowerScale cluster. The following errors may be observed in the /var/log/messages logs on the PowerScale nodes for port 2049, the NFS port: "Error syncache_socket: Socket create failed due to limits or memory shortage" Example:PowerScale-15: 2026-07-28T13:49:14.193127-04:00 <0.7> PowerScale-15(id15) /boot/kernel.amd64/kernel: TCP: [10.10.10.10]:694 to [10.10.10.11]:2049; syncache_socket: Socket create failed due to limits or memory shortage PowerScale-16: 2026-07-28T06:19:00.622215-04:00 <0.7> PowerScale-16(id16) /boot/kernel.amd64/kernel: TCP: [10.10.10.10]:955 to [10.10.10.11]:2049; syncache_socket: Socket create failed due to limits or memory shortage The nodes may start to run out of memory with the following Out of Memory (OOM) errors: 2026-07-16T02:08:15.453461-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: OOM: v_wire_count: 21393044, v_active_count: 3625, v_free_count: 645165, v_inactive_count: 0 events_since_last_log 0s_since_last_log 0 2026-07-16T02:08:15.453511-04:00 <0.4> PowerScale-1(id1) 2026-07-16T02:08:15.453766-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: Malloc Pigs: 2026-07-16T02:08:15.453782-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: Type InUse MemUse Requests 2026-07-16T02:08:15.453797-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: 8kB dinodes 5940489 5141975K 6184573773 2026-07-16T02:08:15.453813-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: crc_vec200 165956 659842K 85328909 2026-07-16T02:08:15.453827-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: isi_hash 209696 604827K 1385867385 2026-07-16T02:08:15.453842-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: iaddr_set 6044672 377792K 279414220502026-07-16T02:08:15.454084-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: UMA Zalloc Pigs: 2026-07-16T02:08:15.454100-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: NAME SIZE LIMIT COUNT MEM USED 2026-07-16T02:08:15.454120-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: IFSINODE 616, 0, 5951191, 3996984K 2026-07-16T02:08:15.454135-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: VNODE 568, 0, 6000000, 3429008K 2026-07-16T02:08:15.454150-04:00 <0.4> PowerScale-1(id1) /boot/kernel.amd64/kernel: VM OBJECT 272, 0, 5998629, 1728396K One may see NFS restarts and NFS coredumps due to the NFS process running out of memory. Stack would be as follows: 2026-07-28T13:48:50.896800-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: [kern_sig.c:4043](pid 46752="nfs")(tid=103791) Stack trace: 2026-07-28T13:48:50.896906-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: Stack: -------------------------------------------------- 2026-07-28T13:48:50.896929-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /lib/libc.so.7:__sys_thr_kill+0xa 2026-07-28T13:48:50.896949-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /lib/libc.so.7:abort+0x49 2026-07-28T13:48:50.896968-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/lw-svcm/nfs.so:$dtrace179609236.NfsAllocateMemoryExplicit+0x11ec 2026-07-28T13:48:50.896984-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/lw-svcm/nfs.so:$dtrace642761941.NfsProtoNfs4ProcReadDir+0x10eb 2026-07-28T13:48:50.896999-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/lw-svcm/nfs.so:$dtrace1357219149.NfsProtoNfs4ProcCompound+0x18a2 2026-07-28T13:48:50.897013-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/lw-svcm/nfs.so:$dtrace1895683854.NfsProtoNfs4Dispatch+0xa31 2026-07-28T13:48:50.897034-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/lw-svcm/nfs.so:NfsExecContextCallback+0x61 2026-07-28T13:48:50.897054-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/liblwsched.so.0:WorkSparkMain+0x4f 2026-07-28T13:48:50.897073-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: /usr/likewise/lib/liblwbase.so.0:SparkMain+0x142 2026-07-28T13:48:50.897088-04:00 <0.5> PowerScale-15(id15) /boot/kernel.amd64/kernel: ----------------------------OR2026-07-26T23:20:31.122983-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: [kern_sig.c:4043](pid 65434="nfs")(tid=103051) Stack trace: 2026-07-26T23:20:31.123076-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: Stack: -------------------------------------------------- 2026-07-26T23:20:31.123112-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: /lib/libc.so.7:__sys_thr_kill+0xa 2026-07-26T23:20:31.123142-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: /lib/libc.so.7:abort+0x49 2026-07-26T23:20:31.123171-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a31274c (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123202-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a3158c4 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123232-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a2e949c (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123259-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a309745 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123284-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a2e42c8 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123312-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a2e3485 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123342-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a2e31f5 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123388-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a2e2f95 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123417-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a393967 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123446-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faf4a450ae6 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123475-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faefadc645d (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123502-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faefadc2898 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123529-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: 0x3faefadd4453 (lookup_symbol: error copying in Ehdr:14) 2026-07-26T23:20:31.123559-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: /lib/libthr.so.3:_pthread_create+0x906 2026-07-26T23:20:31.123590-04:00 <0.5> PowerScale-4(id4) /boot/kernel.amd64/kernel: -------------------------------------------------- 2026-07-26T23:20:31.123619-04:00 <0.6> PowerScale-4(id4) /boot/kernel.amd64/kernel: pid 65434 (nfs), jid 0, uid 0: exited on signal 6 from pid 65434 (self) (core dumped)
OneFS 9.10.1.8 added retry-limit protection to the OneFS zone refresh logic but that fix introduced a counter-reset defect of its own. Before 9.10.1.8, during the NFS exports refresh process, OneFS would throw warnings and simply give up after three attempts if it saw an NFS export with a missing directory. But with the new regression defect, the NFS exports refresh keeps trying infinitely, potentially causing out of memory (OOM) conditions and vnode exhaustion. This is all caused by the constant refresh loops in NFS due to exports unable to find their corresponding directories. Specifically, it is the NFS exports refresh which causes the issue, which occurs when a node is rebooted or when NFS is restarted.
Permanent solution: Upgrade to one of these OneFS versions or later which includes the fix: OneFS 9.10.1.10 PSP-4822 MR:[9.10.1.10_GA-MR][Multiple Userspace and Kernel Fixes](December 2026) OneFS 9.13 (Support has confirmed with ENG that OneFS 9.13 does not include the fix for the original defect which causes this issue; therefore, it is an option for the customer to upgrade to OneFS 9.13 to avoid this issue) In addition, contact Dell PPDM Support to report the issue with PPDM not properly deleting exports as part of the clean-up phase of the backup. Workaround: Until a permanent solution is applied, the following workaround should be used: To prevent the issue, shortly before the planned OneFS 9.10.1.8 upgrade, customers can identify and delete any stale NFS exports (PPDM or otherwise) for which there are missing directories or missing PPDM snapshots. On clusters that have already been upgraded to OneFS 9.10.1.8: Remove all the NFS exports which are missing directories or recreate the missing directories, followed by an NFS restart on all nodes to clear the memory from the NFS processes. Continue to actively monitor for stale NFS exports, especially if the cluster involves PPDM workflow. Note: PowerScale Support has a python script which can identify which exports need to be deleted. Please contact Support for assistance with this option.
Click on a version to see all relevant bugs
Dell Integration
Learn more about where this data comes from
BugZero Plan
Streamline upgrades with automated vendor bug scrubs
BugZero Prevent
Wish you caught this bug sooner? Get proactive today.