OCI Big Data Service provides a managed Hadoop distribution, currently based on Oracle Distribution including Apache Hadoop and Spark. Unlike OCI Data Flow, which runs individual Spark jobs without persistent infrastructure, Big Data Service provisions a persistent cluster with Ambari management, HDFS, Hive, Spark, and the full Hadoop ecosystem. Use it when your workloads need a long-running cluster, HDFS-resident data, Hive tables, or Ambari-managed services that Data Flow cannot provide.
Step 1: IAM Policy
resource "oci_identity_policy" "bds_policy" {
compartment_id = var.compartment_id
name = "big-data-service-policy"
statements = [
"Allow group DataEngineers to manage bds-cluster in compartment id COMPARTMENT_OCID",
"Allow group DataEngineers to use virtual-network-family in compartment id COMPARTMENT_OCID",
"Allow service bds to use virtual-network-family in compartment id COMPARTMENT_OCID",
"Allow service bds to manage instances in compartment id COMPARTMENT_OCID",
"Allow service bds to manage volumes in compartment id COMPARTMENT_OCID"
]
}
Step 2: Provision BDS Cluster
resource "oci_bds_bds_instance" "analytics_cluster" {
compartment_id = var.compartment_id
display_name = "analytics-hadoop-cluster"
cluster_version = "ODH1"
cluster_profile = "HADOOP_EXTENDED"
is_high_availability = true
is_secure = true
network_config {
cidr_block = "10.10.0.0/16"
is_nat_gateway_required = true
}
cluster_details {
bd_cell_version = "COMPATIBLE_WITH_VERSION"
bda_version = "COMPATIBLE_WITH_VERSION"
bdm_version = "COMPATIBLE_WITH_VERSION"
big_data_manager_url = "AUTOGENERATED"
cloudera_manager_url = "AUTOGENERATED"
}
nodes {
shape = "VM.Standard.E4.Flex"
node_type = "MASTER"
subnet_id = var.bds_private_subnet_id
block_volume_size_in_gbs = 150
shape_config {
ocpus = 8
memory_in_gbs = 64
}
}
nodes {
shape = "VM.Standard.E4.Flex"
node_type = "UTILITY"
subnet_id = var.bds_private_subnet_id
block_volume_size_in_gbs = 150
shape_config {
ocpus = 8
memory_in_gbs = 64
}
}
nodes {
shape = "VM.Standard.E4.Flex"
node_type = "WORKER"
subnet_id = var.bds_private_subnet_id
block_volume_size_in_gbs = 1000
number_of_nodes = 4
shape_config {
ocpus = 16
memory_in_gbs = 128
}
}
defined_tags = {
"Operations.Environment" = "production"
"Operations.ManagedBy" = "terraform"
}
}
output "cluster_id" { value = oci_bds_bds_instance.analytics_cluster.id }
Step 3: Auto-Scaling Policy
resource "oci_bds_auto_scaling_configuration" "worker_autoscaler" {
bds_instance_id = oci_bds_bds_instance.analytics_cluster.id
cluster_admin_password = var.cluster_admin_password
display_name = "worker-autoscaler"
is_enabled = true
node_type = "WORKER"
policy_details {
policy_type = "METRIC_BASED_HORIZONTAL_SCALING_POLICY"
scale_up_config {
metric {
metric_type = "CPU_UTILIZATION"
threshold {
duration_in_minutes = 10
operator = "GT"
value = 80
}
}
max_node_count = 12
step_size = 2
memory_step_size = 0
ocpus_step_size = 0
}
scale_down_config {
metric {
metric_type = "CPU_UTILIZATION"
threshold {
duration_in_minutes = 30
operator = "LT"
value = 30
}
}
min_node_count = 4
step_size = 2
memory_step_size = 0
ocpus_step_size = 0
}
}
}
output "autoscaler_id" { value = oci_bds_auto_scaling_configuration.worker_autoscaler.id }
Operational Notes
Choose between Big Data Service and Data Flow based on your data locality requirements. If your Spark jobs read and write to Object Storage and need no HDFS persistence between runs, Data Flow is more cost-effective because the cluster terminates when the job ends. If your jobs use Hive tables, require HDFS replication for intermediate data, or need long-running daemon processes, Big Data Service is the right fit.
Worker auto-scaling uses CPU utilization averaged across all worker nodes. Set the scale-down threshold conservatively at 30 percent and require 30 minutes of sustained low utilization before removing nodes. Hadoop workloads often have long periods of low CPU activity during I/O-bound phases. Aggressive scale-down during these phases causes unnecessary node removal followed immediately by scale-up, which destabilizes the cluster.
Regards,
Osama
#OCI #OracleCloud #BigData #Hadoop #Spark #Terraform #IaC #TechBlog #Oracle #DataEngineering #Analytics #OracleCloudInfrastructure #Hive #HDFS #Ambari #DataPlatform #BigDataAnalytics #AutoScaling #CloudData #DataWarehouse
Leave a comment