Showing posts with label grid. Show all posts
Showing posts with label grid. Show all posts

Monday, October 03, 2011

Oracle Enterprise Manager 12c Grid Control

I've always used Enterprise Manager to some extent. I can remember when it just started out as a Java console program. Clunky, unrefined and not very friendly. Yet, I could see the potential and where Oracle likely wanted to take the product, but wondered as to if they were up to the task. I often thought they would just buy Quest who had an impressive portfolio and was probably (I thought) used most by DBAs, developers, and architects alike. Maybe even Embarcadero, whose DBArtisan tool I actually preferred and found to be a better tool than Quest Toad back then. It is now several years later and how Enterprise Manager has grown up. It is now offered as "DB Console" for single database administration, and as "Grid Control" for multiple database administration and your enterprise-wide monitoring and administration tool.

Grid Control 11g (11.1.0.1) is a web-based tool leveraging WebLogic Server (WLS) 10.3.2 for UI presentation, a database back-end for its storage repository, and agents for data gathering and carrying out monitoring and administration. I'm sure I've been told on many an occasion why each Grid Control will require a specific WLS version but the answer escapes me for now. This however should not be an issue since the WLS should be self-contained (i.e. only for Grid Control) and IMHO installed on the same server as the OMS. If you have multiple OMS for HA, then install a WLS instance on each OMS as well for HA.

Now released is Enterprise Manager 12c (12.1.0.1) Grid Control which will require WLS 10.3.5. This brings a new and improved interface, which also includes cloud computing management features such as charge back and metering.

OEM12.1 - screen shot


My interest in this case are the available upgrade options as shown below:

Two-System Upgrade - Creating a new EM 12.1 infrastructure and switching the new 12.1 agents to point to the new EM 12.1 infrastructure. Requires minimal downtime but new EM infrastructure and you will also have to merge data that was accruing while the upgrade was taking place.

OEM12.1 - 2-system upgrade

One-System Upgrade - This option requires more downtime but no additional infrastructure. Here you are installing the new 12.1 OMS (and 10.3.5 WLS) on the same servers, and then upgrading your OMS to 12.1 (which includes upgrading the repository). Prior to this you would have shutdown your pre-12.1 agents and started up your new 12.1 agents (which were already installed).

OEM12.1 - 1-system upgrade

I prefer the Two-System approach since it has the advantage of minimizing downtime, while providing a fallback option. It also seems a bit less complicated. The previous hardware can be recycled so I don't find that to be an issue (also serves to provide a hardware refresh opportunity). I think most would agree with this being the better of the two options given. To provide a bit more information see the comparison below:

OEM12.1 - upgrade comparison

OEM12.1 - upgrade approaches

OEM12.1 - 2-system upgrade post

OEM12.1 - 1-agent upgrade screen

OEM12.1 - upgrade best practice workflow

Tuesday, July 19, 2011

Oracle 11gR2 (11.2.0.2) Installation - Database software

This is the 2nd in my series of Oracle11gR2 installations, focusing now on the installation of the database or RDBMS software for 11.2.0.2 having already installed the Grid Infrastructure component per my previous post here. If you don't want to install the Grid Infrastructure and use ASM (and Oracle Restart) that is fine, go ahead and skip that first article. In such a case I assume you would be using file systems (perhaps with dNFS which is a post for another time).

Here again, I will be using the command line approach since this is an easy way to have everything scripted and automated (and not require a GUI). I'll show the parameters that need to be adjusted, but if you are not comfortable then I'd suggest doing an initial GUI installation, saving that response file when prompted and then using the saved response file as your gold image for further scripted installations. Note that in the below I use RDBMS_HOME instead of ORACLE_HOME to make the distinction between the actual database home and the grid infrastructure home.

Requirements
See my previous post here

Installation
1. If this is your first installation, then you will want to create the '/etc/oraInst.loc' file, as the root user:

Note: This is optional since you will be prompted at the end of the installation to run $ORACLE_BASE/../oraInventory/orainstRoot.sh which does this for you if this files does not exist.

echo "inventory_loc=/oracle/app/oraInventory
inst_group=dba"
> /etc/oraInst.loc
chown oracle:dba /etc/oraInst.loc
chmod 664 oraInst.loc


2. Edit the DB response file, 'db_inst.rsp' for the values as show below:

ORACLE_HOSTNAME=oradb01
ORACLE_BASE=/oracle/app/oracle
ORACLE_HOME=/oracle/app/product/11.2.0/db_1


3. Run the DB installation (using responsefile and silent installation) as the oracle user:

Note: Ensure you set your DISPLAY environment variable, or you are at run level 5, otherwise you will encounter an error.

./runInstaller -silent -noconfig -responseFile /home/oracle/rsp/db_inst.rsp


4. As the root user run '$RDBMS_HOME/root.sh' when the above completes as prompted.


5. As the oracle user, create an OCM response file. This saves a lot of time later down when you are prompted for those values. Simple run the following and follow the instructions to create and save the response file:

$RDBMS_HOME/OPatch/ocm/bin/emocmrsp


6. Apply the latest OPatch patch (MOS patch ID 6880880), then the latest PSU to this installation as the oracle user. Simple unzip the OPatch patch to the $RDBMS_HOME. For the PSU, unzip to a temporary location, navigate to the patch location and run:

$RDBMS_HOME/OPatch/opatch apply -ocmrf /home/oracle/rsp/ocm.rsp 


7. Apply patch 12431716 (as required by PSU 2) by unzipping to a temporary location, change to the patch directory and as oracle user running: 

$RDBMS_HOME/OPatch/opatch apply -ocmrf /home/oracle/rsp/ocm.rsp


At this point you have fully prepared GI and RDBMS software installations with a ready listener and two disk groups. Now you can create a new database, or migrate an exiting database. I'll leave the new installation to another post in which I'll show how to use DBCA and a template to do it silently, without a GUI.

Thursday, July 07, 2011

Oracle 11gR2 (11.2.0.2) Installation - Grid Infrastructure

I quite often find myself doing Oracle installations and referencing back to the manuals, read me, other blogs, and various other information sources in an attempt to ensure I get everything just perfect. This is regardless of having done it many times and having documented it myself. These days however, I find this a time consuming exercise, especially given my increasing workload so I thought to simply blog about my method which follows below.

The installation below utilizes a command line only method as I often need to be do multiple installations in a standard, and scripted way. The best method I've found for reaching this point is to initially start with a GUI approach, from which the response file is saved which can then be used for further silent, command line driven installations. I'll start with the installation of the Grid Infrastructure component which is used for ASM and Oracle Restart, and continue with a follow-up posting with the RDBMS software installation. In the method that follows I use a simplified setup with only a single group assigned to the oracle user, i.e. 'dba'. Of course you will need to modify the below for your environment.

Requirements
OS: SUSE Linux Enterprise Service 10.3 (64-bit)

Packages (standard installation):
make-3.80-202.2
binutils-2.16.91.0.5-23.34.33
libstdc++33-3.3.3-7.8.1
gcc-4.1.2_20070115-0.29.6
gcc-c++-4.1.2_20070115-0.29.6
glibc-2.4-31.74.1
glibc-devel-2.4-31.74.1
glibc-devel-32bit-2.4-31.74.1
ksh-93s-59.11.35
libaio-0.3.104-14.2
libaio-32bit-0.3.104-14.2
libaio-devel-0.3.104-14.2
libaio-devel-32bit-0.3.104-14.2
libelf-0.8.5-47.2
libgcc-4.1.2_20070115-0.29.6
libstdc++-4.1.2_20070115-0.29.6
libstdc++-devel-4.1.2_20070115-0.29.6
numactl-0.9.6-3.25.28
sysstat-8.0.4-1.7.27
unixODBC-2.2.12
unixODBC-devel-2.2.12
unixODBC-32bit-2.2.12
oracleasmlib-2.0.4-1.SLE10
oracleasm-2.0.5-7.4.50
oracleasm-support-2.1.4-1.SLE10

Configuration files & settings:
/etc/sysctl.conf
# Oracle settings
kernel.core_uses_pid = 1
kernel.suid_dumpable = 1
kernel.msgmni = 2878
# physical RAM size (8 GB) / pagesize (4096)
kernel.shmall = 2097152
# 1/2 of physical RAM typically, 5 GB here
kernel.shmmax = 5368709120
kernel.shmmni = 4096
#kernel.sem = 250 32000 100 128
kernel.sem = 250 32000 100 142
# 512 x processes (for example 6815744 for 13312 processes)
fs.file-max = 6815744
fs.aio-max-nr = 1048576
net.ipv4.ip_local_port_range = 9000 65500
net.ipv4.conf.default.rp_filter = 1
net.core.rmem_default = 262144
net.core.rmem_max = 4194304
net.core.wmem_default = 262144
# For RDS use larger value below
#net.core.wmem_max = 2097152
net.core.wmem_max = 1048576
# group ID of dba group
vm.hugetlb_shm_group = 7001
# number of hugepages (2 MB size) to be used for ASMM
vm.nr_hugepages=2560

/etc/security/limits.conf
# Oracle settings
oracle soft core unlimited
oracle hard core unlimited
oracle soft nproc 131072
oracle hard nproc 131072
oracle soft nofile 131072
oracle hard nofile 131072
oracle soft stack 32768
oracle hard stack 32768
oracle soft memlock 5242880
oracle hard memlock 5242880
root soft core unlimited
root hard core unlimited
root soft nproc 131072
root hard nproc 131072
root soft nofile 131072
root hard nofile 131072
root soft memlock 5242880
root hard memlock 5242880

/etc/pam.d/su
# Oracle setting
session  required       pam_limits.so

/etc/pam.d/xdm
# Oracle setting
session  required       pam_limits.so

/etc/sysconfig/ntp
NTPD_OPTIONS="-x -u ntp"    # add '-x' to existing string

~/.ssh/config
Host *
ForwardX11 no    # Set in the oracle user's file

~/.bashrc
if [ -t 0 ]; then
   stty intr ^C
fi    # Disable terminal output commands (assumes bash shell)

Links
ln -s /bin/fuser /sbin/fuser
ln -s /usr/bin/ssh /usr/local/bin/ssh
ln -s /usr/bin/perl /usr/local/bin/perl
ln -s /usr/bin/scp /usr/local/bin/scp



Installation
1. (optional) Create '/etc/oraInst.loc' file as the root user.

Note: This step is optional as it will be done automatically for first time installations when the oraInst.sh script is run.

echo "inventory_loc=/oracle/app/oraInventory
inst_group=dba"
> /etc/oraInst.loc
chown oracle:dba /etc/oraInst.loc
chmod 664 oraInst.loc


2. Edit GI response file, 'grid.rsp' for the below values:

ORACLE_HOSTNAME=oradb01
ORACLE_BASE=/oracle/app/oracle
ORACLE_HOME=/oracle/app/11.2.0/grid_1


3. As the oracle user, run GI installation (using responsefile 'grid.rsp' and silent installation).

Notes: 
   - Ensure your DISPLAY environment variable is set, or you are at run level 5. Otherwise you will not be able to run the command successfully.
   - You may see some warnings for the installation pertaining to a missing library package compat-libstdc++ which can safely be ignored as for SLES 10.3 and greater libstdc++33 is used instead (reference MOS: 'Requirements for Installing Oracle 11gR2 64-bit (AMD64/EM64T) on SLES 10 (Doc ID 884435.1)').

./runInstaller -silent -noconfig -responseFile /home/oracle/rsp/grid.rsp


4. As the root user, run '/oracle/app/oraInventory/oraInst.sh', '$GI_HOME/root.sh' and then '$GI_HOME/crs/install/roothas.pl' (per log file contents) after the above completes. See below for example run of '$GI_HOME/crs/install/roothas.pl':

/oracle/app/11.2.0/grid_1/perl/bin/perl -I/oracle/app/11.2.0/grid_1/perl/lib -I/oracle/app/11.2.0/grid_1/crs/install /oracle/app/11.2.0/grid_1/crs/install/roothas.pl
Using configuration parameter file: /oracle/app/11.2.0/grid_1/crs/install/crsconfig_params
Creating trace directory
LOCAL ADD MODE
Creating OCR keys for user 'oracle', privgrp 'dba'..
Operation successful.
LOCAL ONLY MODE
Successfully accumulated necessary OCR keys.
Creating OCR keys for user 'root', privgrp 'root'..
Operation successful.
CRS-4664: Node oemgc02 successfully pinned.
Adding daemon to inittab
ACFS-9300: ADVM/ACFS distribution files found.
ACFS-9307: Installing requested ADVM/ACFS software.
ACFS-9308: Loading installed ADVM/ACFS drivers.
ACFS-9321: Creating udev for ADVM/ACFS.
ACFS-9323: Creating module dependencies - this may take some time.
FATAL: Module oracleoks not found.
FATAL: Module oracleadvm not found.
FATAL: Module oracleacfs not found.
ACFS-9327: Verifying ADVM/ACFS devices.
ACFS-9121: Failed to detect control device '/dev/asm/.asm_ctl_spec'.
ACFS-9310: ADVM/ACFS installation failed.
ACFS-9311: not all components were detected after the installation.

oemgc02     2011/06/23 10:02:19     /oracle/app/11.2.0/grid_1/cdata/oemgc02/backup_20110623_100219.olr
Successfully configured Oracle Grid Infrastructure for a Standalone Server

4b. To address the failure of the modules as seen above, reference MOS article "ACFS-9327 ACFS-9121 ACFS-9310: ADVM/ACFS installation failed [ID 1265276.1]" and run the following as the root user:

mkdir -p /lib/modules/`uname -r`/extra/usm
cp -p /lib/modules/2.6.16.60-0.54.5*/extra/usm/* /lib/modules/`uname -r`/extra/usm
/sbin/depmod

4c. OHASD will not starting after reboot on SLES (reference MOS: "OHASD not Starting After Reboot on SLES [ID 1325718.1]"). Run the following to fix this issue:

oradb01:~ # chkconfig --list|grep ohasd
init.ohasd                0:off  1:off  2:off  3:off  4:off  5:off  6:off
ohasd                     0:off  1:off  2:off  3:off  4:off  5:off  6:off
oradb01:~ # chkconfig --list|grep raw
raw                       0:off  1:off  2:off  3:off  4:off  5:off  6:off
oradb01:~ # chkconfig raw on
oradb01:~ # chkconfig ohasd on
oradb01:~ # chkconfig --list|grep raw
raw                       0:off  1:off  2:on   3:on   4:off  5:on   6:off
oradb01:~ # chkconfig --list|grep ohasd
init.ohasd                0:off  1:off  2:off  3:off  4:off  5:off  6:off
ohasd                     0:off  1:off  2:off  3:on   4:off  5:on   6:off

Note: Just turning on OHASD will result in error:

oradb01:~ # chkconfig ohasd on
insserv: Service raw has to be enabled for service oracle_has
insserv: exiting now!
/sbin/insserv failed, exit code 1


5. Apply latest OPatch patch, then latest PSU to GI software installation (Oracle Restart Home). As root user run '$GI_HOME/OPatch/opatch auto /oracle/stage/patches_grid -oh $GI_HOME -ocmrf /home/oracle/rsp/ocm.rsp'. An example below:

oemgc02:/oracle/stage/patches_grid # $ORACLE_HOME/OPatch/opatch auto /oracle/stage/patches_grid -oh $ORACLE_HOME -ocmrf /home/oracle/rsp/ocm.rsp
Executing /usr/bin/perl /oracle/app/11.2.0/grid_1/OPatch/crs/patch112.pl -patchdir /oracle/stage -patchn patches_grid -oh /oracle/app/11.2.0/grid_1 -ocmrf /home/oracle/rsp/ocm.rsp -paramfile /oracle/app/11.2.0/grid_1/crs/install/crsconfig_params
opatch auto log file location is /oracle/app/11.2.0/grid_1/OPatch/crs/../../cfgtoollogs/opatchauto2011-06-23_10-32-00.log
Detected Oracle Restart install
Using configuration parameter file: /oracle/app/11.2.0/grid_1/crs/install/crsconfig_params
Successfully unlock /oracle/app/11.2.0/grid_1
patch /oracle/stage/patches_grid/12311357  apply successful for home  /oracle/app/11.2.0/grid_1
patch /oracle/stage/patches_grid/11724916  apply successful for home  /oracle/app/11.2.0/grid_1
ACFS-9300: ADVM/ACFS distribution files found.
ACFS-9312: Existing ADVM/ACFS installation detected.
ACFS-9314: Removing previous ADVM/ACFS installation.
ACFS-9315: Previous ADVM/ACFS components successfully removed.
ACFS-9307: Installing requested ADVM/ACFS software.
ACFS-9308: Loading installed ADVM/ACFS drivers.
ACFS-9321: Creating udev for ADVM/ACFS.
ACFS-9323: Creating module dependencies - this may take some time.
ACFS-9327: Verifying ADVM/ACFS devices.
ACFS-9309: ADVM/ACFS installation correctness verified.
CRS-4123: Oracle High Availability Services has been started.

5b. Apply patch 12431716, as root user run '$GI_HOME/crs/install/roothas.pl -unlock'. An example below:

oemgc02:/oracle/stage/12431716 # $ORACLE_HOME/crs/install/roothas.pl -unlock
Using configuration parameter file: /oracle/app/11.2.0/grid_1/crs/install/crsconfig_params
Successfully unlock /oracle/app/11.2.0/grid_1

Next, as GI owner run '$GI_HOME/OPatch/opatch apply -oh $GI_HOME -local  -ocmrf /home/oracle/rsp/ocm.rsp'.
Next, as root user run '$GI_HOME/crs/install/roothas.pl -patch'. An example below:

oemgc02:/oracle/stage/12431716 # $ORACLE_HOME/crs/install/roothas.pl -patch
Using configuration parameter file: /oracle/app/11.2.0/grid_1/crs/install/crsconfig_params
ACFS-9300: ADVM/ACFS distribution files found.
ACFS-9312: Existing ADVM/ACFS installation detected.
ACFS-9314: Removing previous ADVM/ACFS installation.
ACFS-9315: Previous ADVM/ACFS components successfully removed.
ACFS-9307: Installing requested ADVM/ACFS software.
ACFS-9308: Loading installed ADVM/ACFS drivers.
ACFS-9321: Creating udev for ADVM/ACFS.
ACFS-9323: Creating module dependencies - this may take some time.
ACFS-9327: Verifying ADVM/ACFS devices.
ACFS-9309: ADVM/ACFS installation correctness verified.
CRS-4123: Oracle High Availability Services has been started.


6. Create listener from GI home (it will automatically get added to CRS by NETCA). By creating the listener before the ASM or database instances, Clusterware will automatically make the dependency relation. As oracle user edit the netca.rsp file and run:

Note: Ensure you have set the DISPLAY environment variable, or you are at run level 5. Otherwise the command will fail.

$GI_HOME/bin/netca -silent -responsefile /home/oracle/resp/netca.rsp


7. Create an ASM instance in silent mode by using ASMCA. The parameters for ASM in 11.2 require a disk group. I like to create a separate disk group (using a 2 GB disk) for this purpose which also serves the purpose of storing the OCR and voting files should the instance be migrated to RAC (since I also like to store those files using a separate disk group). To create this initial disk group, called SYSTEMDG run ASMCA as the oracle (or GI owner) user, then to create the additional disk groups for DATA and FRA, run similar commands using '-createDiskGroup' instead of '-configureASM'. See the examples below:

oracle@oradb01:~/bin> $GI_HOME/bin/asmca -silent -configureASM -sysAsmPassword sysPassw0rd -asmsnmpPassword asmsnmpPassw0rd -diskGroupName SYSTEMDG -diskList 'ORCL:SYSTEMD' -redundancy EXTERNAL -au_size 4 -compatible.asm '11.2.0.2.0' -compatible.rdbms '11.2.0.2.0'

ASM created and started successfully.

Disk Group SYSTEMDG created successfully.

oracle@oradb01:~/bin> $GI_HOME/bin/asmca -silent -createDiskGroup -sysAsmPassword sysPassw0rd -diskGroupName D1_T2_A4 -diskList 'ORCL:D_T2_*' -redundancy EXTERNAL -au_size 4 -compatible.asm '11.2.0.2.0' -compatible.rdbms '11.2.0.2.0'

Disk Group D1_T2_A4 created successfully.

oracle@oradb01:~/bin> $GI_HOME/bin/asmca -silent -createDiskGroup -sysAsmPassword sysPassw0rd -diskGroupName F1_T2_A4 -diskList 'ORCL:F_T2_*' -redundancy EXTERNAL -au_size 4 -compatible.asm '11.2.0.2.0' -compatible.rdbms '11.2.0.2.0'

Disk Group F1_T2_A4 created successfully.


You have now successfully installed Grid Infrastructure 11.2.0.2, applied a PSU and recommended patch, and created an ASM instance (along with a corresponding listener). At this point you are ready to commence the database installation process which I will include in a follow-up blog.

Wednesday, December 08, 2010

Oracle11gR2 RAC: Removing Database Nodes from an Existing Cluster

Phase I - Remove the node from RAC database

Policy-Managed

1. Remove DB Console from the node by running the following from another node in the cluster (as ‘oracle’ user):

$> $RDBMS_HOME/bin/emca -deleteNode db


2. For Policy-Managed database a possible method is to decrease the size of the server pool, and relocate the node to the free pool (assuming all other maximums are met). For relocation to work there should be no active sessions on the associated instance. You can either use the ‘-f’ flag to force relocation, or shutdown the service on the instance prior to running the command:

$> $RDBMS_HOME/bin/srvctl stop instance -d -n
$> $RDBMS_HOME/bin/srvctl relocate server -n -g Free



Admin-Managed

1. For Admin-Managed databases, ensure the instance to be removed is not a PREFERRED or AVAILABLE instance for any Services (i.e. modify Services to exclude the instance to be removed).


2. Remove the instance using ‘$RDBMS_HOME/bin/dbca’, running the command from a node not being removed (as ‘oracle’ user):

$> $RDBMS_HOME/bin/dbca -silent -deleteInstance -nodeList -gdbName -instanceName -sysDBAUserName sys -sysDBAPassword


3. Disable and stop any listeners running on the node (as ‘oracle’ user on any node):

$> $RDBMS_HOME/bin/srvctl disable listener -l -n
$> $RDBMS_HOME/bin/srvctl stop listener -l -n


4. Update the inventory on the node to be removed (run from the node being removed as ‘oracle’ user):

$> $RDBMS_HOME/oui/bin/runInstaller -updateNodeList ORACLE_HOME= “CLUSTER_NODES={[oldnode]}” -local

5. Deinstall the Oracle home, from the node being removed (as ‘oracle’ user):

$> $RDBMS_HOME/deinstall/deinstall -local


6. From any of the existing nodes run the following to update the inventory with the list of the remaining nodes (as ‘oracle’ user):

$> $RDBMS_HOME/oui/bin/runInstaller -updateNodeList ORACLE_HOME= “CLUSTER_NODES={[node1,...nodeX]}"


Phase II - Remove node from Clusterware

7. Check if the node is active and unpinned (as ‘root’ user):

#> $GI_HOME/bin/olsnodes -s -t

Note: The node will only be pinned if using CTSS, or using with database version < style="text-align: justify;">8. Disable Clusterware applications and daemons on the node to be removed. Use the ‘-lastnode’ option when running on the last node in the cluster to be removed (as ‘root’ user):

#> $GI_HOME/crs/install/rootcrs.pl -deconfig -force [-lastnode]


9. From any node not being removed delete Clusterware from the node (as ‘root’ user):

#> $GI_HOME/bin/crsctl delete node -n


10. As ‘grid’ user from node being removed:

$GI_HOME/oui/bin/runInstaller -updateNodeList ORACLE_HOME= "CLUSTER_NODES={[oldnode1,...oldnodeX]}" CRS=TRUE -local

11. Deinstall Clusterware software from the node:

$> $GI_HOME/deinstall/deinstall -local

ISSUE:
‘$GI_HOME/deinstall/deinstall -local’ results in questions about other nodes, and prompts for running scripts against other nodes. DO NOT RUN!! This process deconifigures CRS on all nodes!

ROOT/CAUSE:
Incorrect documentation as a command is missing.

RESOLUTION:
The correct order (as already given in this document) should be as follows:

$GI_HOME/oui/bin/runInstaller -updateNodeList ORACLE_HOME=/oracle/app/11.2.0/grid “CLUSTER_NODES={[oldnode]}” CRS=TRUE -local
$GI_HOME/deinstall/deinstall -local


12. From any existing node to remain, update the Clusterware with the remaining nodes (as ‘grid’ user):

$> $GI_HOME/oui/bin/runInstaller -updateNodeList ORACLE_HOME= “CLUSTER_NODES={[oldnode1,...oldnodeX]}” CRS=TRUE


13. Verify the node has been removed and the remaining nodes are valid:

$GI_HOME/bin/cluvfy stage -post nodedel -n -verbose


14. Remove OCM host/configuration from the MOS portal.

Oracle11gR2 RAC: Adding Database Nodes to an Existing Cluster

Environment:
  • Oracle RAC Database 11.2.0.1
  • Oracle Grid Infrastructure 11.2.0.1
  • Non-GNS
  • OEL 5.5 or SLES 11.1 (both x86_64)

Adding Nodes to a RAC Administrator-Managed Database

Cloning to Extend an Oracle RAC Database (~20 minutes or less depending)

Phase I - Extending Oracle Clusterware to a new cluster node

1. Make physical connections, and install the OS.

Warning: Follow article “11GR2 GRID INFRASTRUCTURE INSTALLATION FAILS WHEN RUNNING ROOT.SH ON NODE 2 OF RAC [ID 1059847.1]” to ensure successful completion of root.sh!! This is also pointed out in the internal Installation Guide.


2. Create Oracle accounts, and setup SSH among the new node and the existing cluster nodes.

Warning: Follow article “How To Configure SSH for a RAC Installation [ID 300548.1]" for the correct procedure for SSH setup! EACH NODE MUST BE VISITED TO ENSURE THEY ARE ADDED TO KNOWN_HOSTS FILE!! This is also mentioned in the internal Installation Guide.


3. Verify the requirements for cluster node addition using the Cluster Verification Utility (CVU). From an existing cluster node (as ‘grid’ user):

$> $GI_HOME/bin/cluvfy stage -post hwos -n -verbose


4. Compare an existing node (reference node) with the new node(s) to be added (as ‘grid’ user):

$> $GI_HOME/bin/cluvfy comp peer -refnode -n -orainv oinstall -osdba dba -verbose


5. Verify the integrity of the cluster and new node by running from an existing cluster node (as ‘grid’ user):

$GI_HOME/bin/cluvfy stage -pre nodeadd -n -fixup -verbose


6. Add the new node by running the following from an existing cluster node (as ‘grid’ user):

a. Not using GNS

$GI_HOME/oui/bin/addNode.sh -silent “CLUSTER_NEW_NODES={}” “CLUSTER_NEW_VIRTUAL_HOSTNAMES={}}”

b. Using GNS

$GI_HOME/oui/bin/addNode.sh -silent “CLUSTER_NEW_NODES={}}”

Run the root scripts when prompted.

POSSIBLE ERROR:
/oracle/app/11.2.0/grid/root.sh
Running Oracle 11g root.sh script...

The following environment variables are set as:
ORACLE_OWNER= grid
ORACLE_HOME= /oracle/app/11.2.0/grid

Enter the full pathname of the local bin directory: [/usr/local/bin]:
Copying dbhome to /usr/local/bin ...
Copying oraenv to /usr/local/bin ...
Copying coraenv to /usr/local/bin ...

Creating /etc/oratab file...
Entries will be added to the /etc/oratab file as needed by
Database Configuration Assistant when a database is created
Finished running generic part of root.sh script.
Now product-specific root actions will be performed.
2010-08-11 16:12:19: Parsing the host name
2010-08-11 16:12:19: Checking for super user privileges
2010-08-11 16:12:19: User has super user privileges
Using configuration parameter file: /oracle/app/11.2.0/grid/crs/install/crsconfig_params
Creating trace directory
-ksh: line 1: /bin/env: not found
/oracle/app/11.2.0/grid/bin/cluutil -sourcefile /etc/oracle/ocr.loc -sourcenode ucstst12 -destfile /oracle/app/11.2.0/grid/srvm/admin/ocrloc.tmp -nodelist ucstst12 ... failed
Unable to copy OCR locations
validateOCR failed for +OCR_VOTE at /oracle/app/11.2.0/grid/crs/install/crsconfig_lib.pm line 7979.

CAUSE:
SSH User Equivalency is not properly setup.

SOLUTION:
[1] Correctly setup SSH user equivalency.

[2] Deconfigure cluster on new node:

#> /oracle/app/11.2.0/grid/crs/install/rootcrs.pl -deconfig -force

[3] Rerun root.sh

POSSIBLE ERROR:
PRCR-1013 : Failed to start resource ora.LISTENER.lsnr
PRCR-1064 : Failed to start resource ora.LISTENER.lsnr on node ucstst13
CRS-2662: Resource 'ora.LISTENER.lsnr' is disabled on server 'ucstst13'

start listener on node=ucstst13 ... failed
Configure Oracle Grid Infrastructure for a Cluster ... failed

CAUSE:
A node is being added that was previously a member of the cluster and Clusterware is aware that the node was a previous member and also that the last listener status was ‘disabled’.

SOLUTION:
Referencing “Bare Metal Restore Procedure for Compute Nodes on an Exadata Environment” (Doc ID 1084360.1), run the following from the node being added as the ‘root’ user to enable and start the local listener (this will complete the operation):

ucstst13:/oracle/app/11.2.0/grid/bin# /oracle/app/11.2.0/grid/bin/srvctl enable listener -l -n
ucstst13:/oracle/app/11.2.0/grid/bin # /oracle/app/11.2.0/grid/bin/srvctl start listener -l
-n


7. Verify that the new node has been added to the cluster (as ‘grid’ user):

$GI_HOME/bin/cluvfy stage -post nodeadd -n -verbose


Phase II - Extending Oracle Database RAC to new cluster node

8. Using the ‘addNode.sh’ script, from an existing node in the cluster as the ‘oracle’ user:

$> $ORACLE_HOME/oui/bin/addNode.sh -silent "CLUSTER_NEW_NODES={newnode1,…newnodeX}"


9. On the new node run the ‘root.sh’ script as the ‘root’ user as prompted.


10. Set ORACLE_HOME and ensure you are using the ‘oracle’ account user. Ensure permissions for Oracle executable are 6751, if not, then as root user:

cd $ORACLE_HOME/bin
chgrp asmadmin oracle
chmod 6751 oracle
ls -l oracle

OR

as 'grid' user: $GI_HOME/bin/setasmgidwrap o=$RDBMS_HOME/bin/oracle


11. On any existing node, run DBCA ($ORACLE_HOME/bin/dbca) to add/create the new instance (as ‘oracle’ user):

$ORACLE_HOME/bin/dbca -silent -addInstance -nodeList -gdbName -instanceName -sysDBAUserName sys -sysDBAPassword

NOTE: Ensure the command is run from an existing node with the same or less memory than the new nodes otherwise the command will fail due to insufficient memory to support the instance. Also ensure that the log file is checked for actual success since it can differ from what is displayed at the screen.

POSSIBLE ERROR:
DBCA logs with error below (screen indicates success):
Adding instance
DBCA_PROGRESS : 1%
DBCA_PROGRESS : 2%
DBCA_PROGRESS : 6%
DBCA_PROGRESS : 13%
DBCA_PROGRESS : 20%
DBCA_PROGRESS : 26%
DBCA_PROGRESS : 33%
DBCA_PROGRESS : 40%
DBCA_PROGRESS : 46%
DBCA_PROGRESS : 53%
DBCA_PROGRESS : 66%
Completing instance management.
DBCA_PROGRESS : 76%
PRCR-1013 : Failed to start resource ora.racdb.db
PRCR-1064 : Failed to start resource ora.racdb.db on node ucstst13
ORA-01031: insufficient privileges
ORA-01034: ORACLE not available
ORA-27101: shared memory realm does not exist
Linux-x86_64 Error: 2: No such file or directory
Process ID: 0
Session ID: 0 Serial number: 0

ORA-01031: insufficient privileges
ORA-01031: insufficient privileges
CRS-2674: Start of 'ora.racdb.db' on 'ucstst13' failed
ORA-01034: ORACLE not available
ORA-27101: shared memory realm does not exist
Linux-x86_64 Error: 2: No such file or directory
Process ID: 0
Session ID: 0 Serial number: 0

ORA-01031: insufficient privileges
ORA-01031: insufficient privileges

DBCA_PROGRESS : 100%

CAUSE:
The permissions and/or ownership on the ‘$RDBMS_HOME/bin/oracle’ binary are incorrect. References: “ORA-15183 Unable to Create Database on Server using 11.2 ASM and Grid Infrastructure [ID 1054033.1]”, “Incorrect Ownership and Permission after Relinking or Patching 11gR2 Grid Infrastructure [ID 1083982.1]”.

SOLUTION:
As root user:
cd $ORACLE_HOME/bin
chgrp asmadmin oracle
chmod 6751 oracle
ls -l oracle

Ensure the ownership and permission are now like:
-rwsr-s--x 1 oratest asmadmin

Warning: Whenever a patch is applied to the database ORACLE_HOME, please ensure the above ownership and permission are corrected after the patch.


12. Following the above, the EM DB Console may not be configured correctly for the new nodes. The GUI does not show the nodes, however the EM agent does report them to the OMS agent via the command line:

**************** Current Configuration ****************
INSTANCE NODE DBCONTROL_UPLOAD_HOST
---------- ---------- ---------------------

racdb racnode10 racnode10.mydomain.com
racdb racnode11 racnode10.mydomain.com
racdb racnode12 racnode10.mydomain.com
racdb racnode13 racnode10.mydomain.com
racdb racnode14 racnode10.mydomain.com

Also, the listener is not properly configured as starting from the GI_HOME so it will appear incorrectly as down in DB Console. To correctly configure the instance in DB Console:

a. Delete the instance from EM DB Console:

$RDBMS_HOME/bin/emca -deleteInst db

b. Create the EM Agent directories on the new node for all the nodes in the cluster (including the new node) as follows:

mkdir -p $RDBMS_HOME/racnode10_racdb/sysman/config
mkdir -p $RDBMS_HOME/racnode11_racdb/sysman/config
mkdir -p $RDBMS_HOME/racnode12_racdb/sysman/config
mkdir -p $RDBMS_HOME/racnode13_racdb/sysman/config
mkdir -p $RDBMS_HOME/racnode10_racdb/sysman/emd
mkdir -p $RDBMS_HOME/racnode11_racdb/sysman/emd
mkdir -p $RDBMS_HOME/racnode12_racdb/sysman/emd
mkdir -p $RDBMS_HOME/racnode13_racdb/sysman/emd

c. Re-add the instance:

$RDBMS_HOME/emca -addInst db

To address the listener incorrectly showing as down in DB Console:

a. On each node as the oracle user, edit the file ‘$RDBMS_HOME/_/sysman/config/targets.xml’.

b. Change the "ListenerOraDir" entry for the node to match the ‘$GI_HOME/network/admin’ location.

c. Save the file and restart the DB Console.


13. Verify the administrator privileges on the new node by running on existing node:

$ORACLE_HOME/bin/cluvfy comp admprv -o db_config -d -n -verbose


14. For an Admin-Managed Cluster, add the new instances to Services, or create additional Services. For a Policy-Managed Cluster, verify the instance has been added to an existing server pool.

15. Setup OCM in the cloned homes (both GI and RDBMS):

a. Delete all subdirectories to remove previously configured host:

$> rm -rf $ORACLE_HOME/ccr/hosts/*

b. Move (do not copy) from the OH the file as shown below:

$> mv $ORACLE_HOME/ccr/inventory/core.jar $ORACLE_HOME/ccr/inventory/pending/core.jar

c. Configure OCM for the cloned home on the new node:

$> $ORACLE_HOME/ccr/bin/configCCR -a



Adding Nodes to a RAC Policy-Managed Database
When adding a node in a cluster running as a Policy-Managed Database, Oracle Clusterware tries to start the new instance before the cloning procedure completes. The following steps should be used to add the node:

1. Run the ‘addNode.sh’ script as the ‘grid’ user for the Oracle Grid Infrastructure for a cluster to add the new node (similar to step 6 above). DO NOT run the root scripts when prompted; you will run them later.


2. Install the Oracle RAC database software using a software-only installation, clone or ‘addNode.sh’ script (same as step 8 above) method. Ensure Oracle is linked with the Oracle RAC option if using the software-only installation.


3. Complete the root script actions for the Database home, similar to steps 9 - 10 above.


4. Complete the root scripts action for the Oracle Clusterware home and then finish the installation. Similar to as mentioned in step 6 above.


5. Verify that EM DB Console is fully operational, similar to step 12 above.


6. Complete the configuration for OCM, similar to step 15 above.