Thursday, January 16, 2014

How to map a snapshot of a LUN to another server for backup

Issue:
You would like to map the Data OnTap snapshot of a LUN to another server as a LUN to be used for daily/weekly backups to tape over Fiber Channel or iSCSI.  This scenario is commonly referred to as off-host backup.

Solution:
To obtain a consistent snapshot you must use the SnapDrive or SnapManager products from Netapp to create the snapshot of the desired LUN.  Then on the off-host backup server you can use SnapDrive to mount the snapshot as a LUN.

SnapDrive mounts the snapshot of the LUN as a Read/Write Clone LUN, not as Read Only, but changes to the snapshot LUN are deleted once the LUN is deleted.

Workaround:
As an alternative you can create an inconsistent snapshot on the Netapp Controller use that snapshot as the foundation of a LUN Clone.

Here are the commands to run on the Netapp Controller to take any existing snapshot and create a LUN Clone

1)      Optional, create a new snap to be used as the basis of the LUN clone
snap create volume snap_name

2)      Create the lun clone, in the same volume as existing lun, using existing snap
lun clone create /vol/volume/clone01 -b /vol/volume/lun01 snap_name

a.       If you get the following error lun clone create: No space left on device use the following syntax
lun clone create /vol/volume/clone01 -o noreserve -b /vol/volume/lun01 snap_name

3)      Mount the clone to the off-host backup server.
lun map /vol/volume/clone01 backup_igroup #

backup_igroup = the initiator group associated with the backup server.
# = the LUN ID to mount the LUN as

a.       You could use SnapDrive to mount the LUN to the off-host backup server, as an alternative.
4)      On the backup server rescan your drives
5)      Bring the new volume on-line
6)      Perform the backup
7)      Take the volume off-line
8)      Remove the clone from the off-host backup server
lun unmap /vol/volume/clone01 backup_igroup

backup_igroup = the initiator group associated with the backup server

9)      Remove the clone
lun destroy /vol/volume/clone01

How to fix WAFL hung in SK process

Issue:
You receive the following error message during a failover of Netapp controllers.

WAFL hung in SK process idle_thread0 on release NetApp Release 8.0.1

Cause:
The most likely cause is Netapp bug #455404.  In general this bug is caused because WAFL runs out of buffers due to a memory leak.

Solution:
Contact Netapp support for a core dump analysis, most likely they will find WAFL ran out of buffers.

If your Netapp controllers are configured in a clustered pair you should perform a giveback ASAP.  This should be done just in case the WAFL bug hits the other controller, you will be in a fault tolerant state.

More Information:
Log onto the Netapp Now site and use the Bug Tool to review bug #455404 for a list of Data Ontap versions in which this bug was fixed.  In general these are.

  • Data ONTAP 7.3.6RC1
  • Data ONTAP 7.3.6 (GA)
  • Data ONTAP 8.0.2 (GA)
You may want to schedule an upgrade to one of the fixed Data Ontap versions if needed, contact Netapp support for more information.

WAFL is an acronym for Write Anywhere File Layout and is a type of file system used by Data Ontap.

SK is a process used by WAFL to do cleanup of the WAFL file system.

Thursday, January 9, 2014

Updating SP firmware

Updating SP firmware

If your storage system includes a Service Processor (SP), you must verify that it is running the correct firmware version and update the firmware if it is not the correct version.

Before you begin

The reversion or downgrade process should be complete and the storage system should be running the target release.

About this task

Data ONTAP software images include firmware for SP modules. If the firmware version on your SP module is outdated, you must update it before returning the reverted or downgraded system to production status.

Steps

  1. Go to the system firmware information on the NetApp Support Site and determine the most recent firmware version for your SP module.
  2. Enter the following command at the storage system CLI to determine the SP firmware version:
    sp status
    You see output similar to the following:
    Service Processor Status: Online
    Firmware Version:   1.2
    ...
    
    If the SP firmware version in the command output is earlier than the most recent version on the NetApp Support Site, you must update your disk shelf firmware manually.
  3. Click the SP_FW.zip link to download the file from the NetApp Support Site to your HTTP server.
  4. At the storage system prompt, enter the following command:
    software update http://web_server/SP_FW.zip -f
  5. When the software update command is finished, enter the following command at the storage system prompt:
    sp update
  6. When the system prompts you to update the SP, enter y to continue.
    The SP is updated and you are prompted to reboot the SP. Wait approximately 60 seconds to allow the SP to reboot.
  7. Verify that the SP firmware has been updated by entering the following command:
    sp status
  8. If the system is a partner in an HA pair, repeat Steps 4 through 7 on the partner system.

Result

If your console connection is not through the SP, the connection remains active during the SP reboot.
If your console connection is through the SP, you lose your console connection to the storage system. In approximately one minute, the SP reboots and automatically reestablishes the connection.

Wednesday, January 8, 2014

How to troubleshoot takeover of partner is disabled due to unsynchronized log

Issue:
You receive one or more of these messages from your Netapp syslog server.

2011-10-03 09:49:21 Kernel.Debug netapp2 Nov 3 09:48:21 [iwarp-vfiler@netapp2: ctrl.rdma.failConnect:debug]: Failed to connect.

2011-10-03 09:58:20 Kernel.Notice netapp2 Nov 3 09:57:20 [netapp2: cf.fsm.partnerNotResponding:notice]: Failover monitor: partner not responding

2011-10-03 10:01:08 Kernel.Warning netapp2 Nov 3 09:00:08 [netapp2: cf.takeover.disabled:warning]: Controller Failover is licensed but takeover of partner is disabled due to reason : unsynchronized log.

Filer View may report the following
Controller failover of hostname is not possible: unsynchronized log

Cause:
Typical causes of the interconnect going up and down are
  • Loose connections on the interconnect cable.
  • Bad interconnect cable
  • Bad internal connection or port on the Netapp Controller.

Solution:
In general call Netapp Support to help diagnose and troubleshoot these error messages.  If the problem requires parts replacement you really should have Netapp support do the replacement.

More Information:
Troubleshooting steps you can take on your own.

1)  Run the following command
cf status

a.  If you receive a message similar to this, call Netapp Support asap.

netapp1 is up, takeover disabled because of reason (unsynchronized log)
netapp2 has disabled takeover by netapp1 (unsynchronized log)
Interconnect status: down.

2)  Check the interconnect cables between the two Netapp Controllers
a.  The type of connect may vary by Netapp Controller model, in general
  i.  Check for loose connections
  ii.  Check for amber lights (normal status is a green light)
b.  If the interconnect is down you can disconnect and reconnect the interconnect cable(s) as needed in an attempt to bring the connection back up.
3)  Depending on your controller type this may or may not apply.
a.  If your controller uses two cables/ports (example: FAS 3240 uses ports c0a and c0b), there should be 4 lighted arrows corresponding to the ports as follows.
  i.  Top port: corresponds to the 2 lighted arrows pointing up.
  ii.  Bottom port: corresponds to the 2 lighted arrows pointing down.
b.  So the following should hold true for two port interconnect configurations
  i.  Two up arrows or two down arrows (colored amber) indicate a bad cable or individual port.
·  It also follows you would see corresponding Link down messages in the syslog similar to the following for that port.

netif.linkDown:info Ethernet c0b: Link down, check cable.

  ii.  One up and one down arrow (colored amber) indicates a bad internal connection on one of the Netapp Controllers. 
·  It also follows you should not see Link down messages in the syslog, but should see link messages for up and down when you can disconnect and reconnect the cable from the port.

Effects or ramifications of the interconnect being down
The Netapp Controller interconnect is used in the failover (takeover/giveback) process, while the controller interconnect is down this process will not work.  If a takeover is initiated the controller will crash, requiring a manual restart of the logical controller on either physical Netapp Controllers.

The Netapp Controller interconnect is also used to pass Misconfigured Partner Path (aka: secondary path) IOs back to the Primary Controller for the given volume/lun.  When the controller interconnect is down this functionality is not available, so IOs over secondary paths will fail. 

If this happens you will need to correct the server configuration to use the primary paths or disable the secondary paths as needed on your servers until the controller interconnect is back up.

If you are using the Microsoft DSM your servers should detect the secondary paths as unavailable automatically.

The Netapp DSM may not detect the interconnect as being down, and may instead report an error with the LUN.

Definitions
RDMA (Remote Direct Memory Access): is used for high performance data transfer between controllers configured in an High Availability (HA) pair and is done over the controller interconnect.  This technology is part of the controller to controller communications providing the takeover/giveback and secondary path functionality of a HA pair.


Known Bug
Bug ID 489576
Title FAS/V3200 series storage failover disabled due to "unsynchronized log"

Description
After some period of uptime with storage failover enabled, a FAS/V3200 series system may suddenly encounter an "unsynchronized log" condition that it cannot recover from. This will result in the loss of high-availability failover and SFO will be disabled.

This event is triggered by a timing error in the onboard 10Gb controller firmware that is responsible for storage cluster communications on this platform. This problem may occur with any FAS/V3200A or FAS/V3200AE configuration running 7.3.5 or 8.0.1 versions of Data ONTAP.

Workaround
A reboot will clear the condition temporarily.

The bug is fixed in the following versions
Data ONTAP 7.3.6RC1 (First Fixed) - Fixed
Data ONTAP 7.3.6 (GA) - Fixed
Data ONTAP 8.0.2 (GA) - Fixed
Data ONTAP 8.1RC1 (RC) - Fixed

Tuesday, January 7, 2014

Netapp NFS Exportfs CLI Configuration Guide

Netapp CLI NFS Config Guide

A quick and simple Netapp NFS configuration guide with commands and options to help explain and remove the mysteries. Netapps provide highly dependable NFS services, as the name implies, it’s a network appliance. You really can just turn it on and not worry much about outages. Unless someone trips on a power chord or two. Below is a compilation of exporfs and exports configuration options commonly used.
Rules for exporting Resources
• Specify complete pathname, meaning the path must begin with a /vol prefix
• You cannot export /vol, which is not a pathname to a file, directory or volume. Export each volume separately
• When export a resource to multiple targets, separate the target names with a colon (:) Resolve hostnames using DNS, NIS or /etc/hosts per order in /etc/nssswitch.conf
Examples to export resources with NFS on the CLI
> exportfs -a
> exportfs -o rw=host1:host2 /vol/volxyz
Exportable resources are by Volume Directory/Qtree File. Target examples from /etc/exports Host – use name of IP address
/vol/vol0/home -rw=myhost
/vol/vol0/home -root=myhost,-rw=hishost,therehost
Netgroup – use the NIS group name – Although I love NIS, NIS is rare nowadays. But the Netapp supports NIS. Don’t think NIS+.
/vol/vol0/home -rw=the-nisgroup
Subnet – specify the subnet address
/vol/vol0/home -rw=”192.168.100.0/24″
DNS – use DNS subdomain
/vol/vol0/home -rw=”.sap.dev.mydomain.com”
Displays all current export in memory
> exportfs
To export all file system paths specified in the /etc/exports file
> exportfs -a
Adds exports to the /etc/exports file and in memory.  Default export options are always “rw” (all hosts) and  security set to “sec=sys”
> exportfs -p [options] path
> exports -p rw=hostxyz /vol/vol2/sap
Exports a file system path temporarily without adding a corresponding entry to the /etc/exports file, handy short term method.
> exporfs -i -o ro=hostB /vol/vol1/lun2
Reloads the exports from /etc/exports files
> exportfs -r
Unexports all exports defined in the /etc/exports file
> exportfs -uav
Unexports a specific export
> exportfs -u /vol/vol2/homes
Unexports an export and removes it from /etc/exports file. This one is handy.
> exportfs -z /vol/vol0/home
To verify the actual path to which a volume is exported
> exportfs -s /vol/vol2/vms-data
To display list of clients mounting from the storage system
> showmount -a filerabc
 To display list of exported resources on the storage system
>showmount -e filerabc
 To check NFS target to access cache
> exportfs -c clientaddr path [accesstype] [securitytype]
> exportfs -c host1 /vol/vol2 rw
To remove access cache entries
> exportfs -f [path]
Flush the access cache.
> exportfs -f
 Access restrictions that specify what operations an NFS client can perform on a resource
• Default is read-write (rw) and UNIX Auth_SYS (sys) security
• “ro” option provides read-ony access to all hosts
• “ro=” option provides read-only access to specified hosts
• “rw=” option provides read-write access to specified hosts
• “root=” option specifies that root on the target has root permissions