DecisionTreeClassificationModel (Spark 3.4.1 JavaDoc)

Object
- org.apache.spark.ml.PipelineStage
- - org.apache.spark.ml.Transformer
  - - org.apache.spark.ml.Model<M>
    - - org.apache.spark.ml.PredictionModel<FeaturesType,M>
      - org.apache.spark.ml.classification.ClassificationModel<FeaturesType,M>
        
        org.apache.spark.ml.classification.ProbabilisticClassificationModel<Vector,DecisionTreeClassificationModel>
        
        org.apache.spark.ml.classification.DecisionTreeClassificationModel

All Implemented Interfaces:

java.io.Serializable, org.apache.spark.internal.Logging, ClassifierParams, ProbabilisticClassifierParams, Params, HasCheckpointInterval, HasFeaturesCol, HasLabelCol, HasPredictionCol, HasProbabilityCol, HasRawPredictionCol, HasSeed, HasThresholds, HasWeightCol, PredictorParams, DecisionTreeClassifierParams, DecisionTreeModel, DecisionTreeParams, TreeClassifierParams, Identifiable, MLWritable
```
public class DecisionTreeClassificationModel
extends ProbabilisticClassificationModel<Vector,DecisionTreeClassificationModel>
implements DecisionTreeModel, DecisionTreeClassifierParams, MLWritable, scala.Serializable
```
Decision tree model (http://en.wikipedia.org/wiki/Decision_tree_learning) for classification. It supports both binary and multiclass labels, as well as both continuous and categorical features.

See Also:

Serialized Form

Nested Class Summary
- Nested classes/interfaces inherited from interface org.apache.spark.internal.Logging
  org.apache.spark.internal.Logging.SparkShellLoggingFilter

Method Summary

All Methods Static Methods Instance Methods Concrete Methods
Modifier and Type	Method and Description
`BooleanParam`	`cacheNodeIds()` If false, the algorithm will pass trees to executors to match instances with nodes.
`IntParam`	`checkpointInterval()` Param for set checkpoint interval (>= 1) or disable checkpoint (-1).
`DecisionTreeClassificationModel`	`copy(ParamMap extra)` Creates a copy of this instance with the same UID and some extra params.
`int`	`depth()` Depth of the tree.
`Vector`	`featureImportances()`
`Param<String>`	`impurity()` Criterion used for information gain calculation (case-insensitive).
`Param<String>`	`leafCol()` Leaf indices column name.
`static DecisionTreeClassificationModel`	`load(String path)`
`IntParam`	`maxBins()` Maximum number of bins used for discretizing continuous features and for choosing how to split on features at each node.
`IntParam`	`maxDepth()` Maximum depth of the tree (nonnegative).
`IntParam`	`maxMemoryInMB()` Maximum memory in MB allocated to histogram aggregation.
`DoubleParam`	`minInfoGain()` Minimum information gain for a split to be considered at a tree node.
`IntParam`	`minInstancesPerNode()` Minimum number of instances each child must have after split.
`DoubleParam`	`minWeightFractionPerNode()` Minimum fraction of the weighted sample count that each child must have after split.
`int`	`numClasses()` Number of classes (values which the label can take).
`int`	`numFeatures()` Returns the number of features the model was trained on.
`double`	`predict(Vector features)` Predict label for the given features.
`Vector`	`predictRaw(Vector features)` Raw prediction for each possible label.
`static MLReader<DecisionTreeClassificationModel>`	`read()`
`Node`	`rootNode()` Root of the decision tree
`LongParam`	`seed()` Param for random seed.
`String`	`toString()` Summary of the model
`Dataset<Row>`	`transform(Dataset<?> dataset)` Transforms dataset by reading from `featuresCol`, and appending new columns as specified by parameters: - predicted labels as `predictionCol` of type `Double` - raw predictions (confidences) as `rawPredictionCol` of type `Vector` - probability of each class as `probabilityCol` of type `Vector`.
`StructType`	`transformSchema(StructType schema)` Check transform validity and derive the output schema from the input schema.
`String`	`uid()` An immutable unique ID for the object and its derivatives.
`Param<String>`	`weightCol()` Param for weight column name.
`MLWriter`	`write()` Returns an `MLWriter` instance for this ML instance.

Methods inherited from class org.apache.spark.ml.classification.ProbabilisticClassificationModel
normalizeToProbabilitiesInPlace, predictProbability, probabilityCol, setProbabilityCol, setThresholds, thresholds

Methods inherited from class org.apache.spark.ml.classification.ClassificationModel
rawPredictionCol, setRawPredictionCol, transformImpl

Methods inherited from class org.apache.spark.ml.PredictionModel
featuresCol, labelCol, predictionCol, setFeaturesCol, setPredictionCol

Methods inherited from class org.apache.spark.ml.Model
hasParent, parent, setParent

Methods inherited from class org.apache.spark.ml.Transformer
transform, transform, transform

Methods inherited from class org.apache.spark.ml.PipelineStage
params

Methods inherited from class Object
equals, getClass, hashCode, notify, notifyAll, wait, wait, wait

Methods inherited from interface org.apache.spark.ml.tree.DecisionTreeModel
getLeafField, leafIterator, maxSplitFeatureIndex, numNodes, predictLeaf, toDebugString

Methods inherited from interface org.apache.spark.ml.tree.DecisionTreeClassifierParams
validateAndTransformSchema

Methods inherited from interface org.apache.spark.ml.tree.DecisionTreeParams
getCacheNodeIds, getLeafCol, getMaxBins, getMaxDepth, getMaxMemoryInMB, getMinInfoGain, getMinInstancesPerNode, getMinWeightFractionPerNode, getOldStrategy, setLeafCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasCheckpointInterval
getCheckpointInterval

Methods inherited from interface org.apache.spark.ml.param.shared.HasSeed
getSeed

Methods inherited from interface org.apache.spark.ml.param.shared.HasWeightCol
getWeightCol

Methods inherited from interface org.apache.spark.ml.tree.TreeClassifierParams
getImpurity, getOldImpurity

Methods inherited from interface org.apache.spark.ml.param.shared.HasLabelCol
getLabelCol, labelCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasFeaturesCol
featuresCol, getFeaturesCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasPredictionCol
getPredictionCol, predictionCol

Methods inherited from interface org.apache.spark.ml.param.Params
clear, copyValues, defaultCopy, defaultParamMap, explainParam, explainParams, extractParamMap, extractParamMap, get, getDefault, getOrDefault, getParam, hasDefault, hasParam, isDefined, isSet, onParamChange, paramMap, params, set, set, set, setDefault, setDefault, shouldOwn

Methods inherited from interface org.apache.spark.ml.param.shared.HasRawPredictionCol
getRawPredictionCol, rawPredictionCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasProbabilityCol
getProbabilityCol, probabilityCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasThresholds
getThresholds, thresholds

Methods inherited from interface org.apache.spark.ml.util.MLWritable
save

Methods inherited from interface org.apache.spark.internal.Logging
$init$, initializeForcefully, initializeLogIfNecessary, initializeLogIfNecessary, initializeLogIfNecessary$default$2, initLock, isTraceEnabled, log, logDebug, logDebug, logError, logError, logInfo, logInfo, logName, logTrace, logTrace, logWarning, logWarning, org$apache$spark$internal$Logging$$log__$eq, org$apache$spark$internal$Logging$$log_, uninitialize

- Method Detail
  - read
```
public static MLReader<DecisionTreeClassificationModel> read()
```
  - load
```
public static DecisionTreeClassificationModel load(String path)
```
  - impurity
```
public final Param<String> impurity()
```
    Description copied from interface: TreeClassifierParams
    
    Criterion used for information gain calculation (case-insensitive). This impurity type is used in DecisionTreeClassifier and RandomForestClassifier, Supported: "entropy" and "gini". (default = gini)
    
    Specified by:
    
    impurity in interface TreeClassifierParams
    
    Returns:
    
    (undocumented)
  - leafCol
```
public final Param<String> leafCol()
```
    Description copied from interface: DecisionTreeParams
    
    Leaf indices column name. Predicted leaf index of each instance in each tree by preorder. (default = "")
    
    Specified by:
    
    leafCol in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - maxDepth
```
public final IntParam maxDepth()
```
    Description copied from interface: DecisionTreeParams
    
    Maximum depth of the tree (nonnegative). E.g., depth 0 means 1 leaf node; depth 1 means 1 internal node + 2 leaf nodes. (default = 5)
    
    Specified by:
    
    maxDepth in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - maxBins
```
public final IntParam maxBins()
```
    Description copied from interface: DecisionTreeParams
    
    Maximum number of bins used for discretizing continuous features and for choosing how to split on features at each node. More bins give higher granularity. Must be at least 2 and at least number of categories in any categorical feature. (default = 32)
    
    Specified by:
    
    maxBins in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - minInstancesPerNode
```
public final IntParam minInstancesPerNode()
```
    Description copied from interface: DecisionTreeParams
    
    Minimum number of instances each child must have after split. If a split causes the left or right child to have fewer than minInstancesPerNode, the split will be discarded as invalid. Must be at least 1. (default = 1)
    
    Specified by:
    
    minInstancesPerNode in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - minWeightFractionPerNode
```
public final DoubleParam minWeightFractionPerNode()
```
    Description copied from interface: DecisionTreeParams
    
    Minimum fraction of the weighted sample count that each child must have after split. If a split causes the fraction of the total weight in the left or right child to be less than minWeightFractionPerNode, the split will be discarded as invalid. Should be in the interval [0.0, 0.5). (default = 0.0)
    
    Specified by:
    
    minWeightFractionPerNode in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - minInfoGain
```
public final DoubleParam minInfoGain()
```
    Description copied from interface: DecisionTreeParams
    
    Minimum information gain for a split to be considered at a tree node. Should be at least 0.0. (default = 0.0)
    
    Specified by:
    
    minInfoGain in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - maxMemoryInMB
```
public final IntParam maxMemoryInMB()
```
    Description copied from interface: DecisionTreeParams
    
    Maximum memory in MB allocated to histogram aggregation. If too small, then 1 node will be split per iteration, and its aggregates may exceed this size. (default = 256 MB)
    
    Specified by:
    
    maxMemoryInMB in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - cacheNodeIds
```
public final BooleanParam cacheNodeIds()
```
    Description copied from interface: DecisionTreeParams
    
    If false, the algorithm will pass trees to executors to match instances with nodes. If true, the algorithm will cache node IDs for each instance. Caching can speed up training of deeper trees. Users can set how often should the cache be checkpointed or disable it by setting checkpointInterval. (default = false)
    
    Specified by:
    
    cacheNodeIds in interface DecisionTreeParams
    
    Returns:
    
    (undocumented)
  - weightCol
```
public final Param<String> weightCol()
```
    Description copied from interface: HasWeightCol
    
    Param for weight column name. If this is not set or empty, we treat all instance weights as 1.0.
    
    Specified by:
    
    weightCol in interface HasWeightCol
    
    Returns:
    
    (undocumented)
  - seed
```
public final LongParam seed()
```
    Description copied from interface: HasSeed
    
    Param for random seed.
    
    Specified by:
    
    seed in interface HasSeed
    
    Returns:
    
    (undocumented)
  - checkpointInterval
```
public final IntParam checkpointInterval()
```
    Description copied from interface: HasCheckpointInterval
    
    Param for set checkpoint interval (>= 1) or disable checkpoint (-1). E.g. 10 means that the cache will get checkpointed every 10 iterations. Note: this setting will be ignored if the checkpoint directory is not set in the SparkContext.
    
    Specified by:
    
    checkpointInterval in interface HasCheckpointInterval
    
    Returns:
    
    (undocumented)
  - depth
```
public int depth()
```
    Description copied from interface: DecisionTreeModel
    
    Depth of the tree. E.g.: Depth 0 means 1 leaf node. Depth 1 means 1 internal node and 2 leaf nodes.
    
    Specified by:
    
    depth in interface DecisionTreeModel
    
    Returns:
    
    (undocumented)
  - uid
```
public String uid()
```
    Description copied from interface: Identifiable
    
    An immutable unique ID for the object and its derivatives.
    
    Specified by:
    
    uid in interface Identifiable
    
    Returns:
    
    (undocumented)
  - rootNode
```
public Node rootNode()
```
    Description copied from interface: DecisionTreeModel
    
    Root of the decision tree
    
    Specified by:
    
    rootNode in interface DecisionTreeModel
  - numFeatures
```
public int numFeatures()
```
    Description copied from class: PredictionModel
    
    Returns the number of features the model was trained on. If unknown, returns -1
    
    Overrides:
    
    numFeatures in class PredictionModel<Vector,DecisionTreeClassificationModel>
  - numClasses
```
public int numClasses()
```
    Description copied from class: ClassificationModel
    
    Number of classes (values which the label can take).
    
    Specified by:
    
    numClasses in class ClassificationModel<Vector,DecisionTreeClassificationModel>
  - predict
```
public double predict(Vector features)
```
    Description copied from class: ClassificationModel
    
    Predict label for the given features. This method is used to implement transform() and output predictionCol.
    This default implementation for classification predicts the index of the maximum value from predictRaw().
    
    Overrides:
    
    predict in class ClassificationModel<Vector,DecisionTreeClassificationModel>
    
    Parameters:
    
    features - (undocumented)
    
    Returns:
    
    (undocumented)
  - transformSchema
```
public StructType transformSchema(StructType schema)
```
    Description copied from class: PipelineStage
    
    Check transform validity and derive the output schema from the input schema.
    We check validity for interactions between parameters during transformSchema and raise an exception if any parameter value is invalid. Parameter value checks which do not depend on other parameters are handled by Param.validate().
    Typical implementation should first conduct verification on schema change and parameter validity, including complex parameter interaction checks.
    
    Overrides:
    
    transformSchema in class ProbabilisticClassificationModel<Vector,DecisionTreeClassificationModel>
    
    Parameters:
    
    schema - (undocumented)
    
    Returns:
    
    (undocumented)
  - transform
```
public Dataset<Row> transform(Dataset<?> dataset)
```
    Description copied from class: ProbabilisticClassificationModel
    
    Transforms dataset by reading from featuresCol, and appending new columns as specified by parameters: - predicted labels as predictionCol of type Double - raw predictions (confidences) as rawPredictionCol of type Vector - probability of each class as probabilityCol of type Vector.
    
    Overrides:
    
    transform in class ProbabilisticClassificationModel<Vector,DecisionTreeClassificationModel>
    
    Parameters:
    
    dataset - input dataset
    
    Returns:
    
    transformed dataset
  - predictRaw
```
public Vector predictRaw(Vector features)
```
    Description copied from class: ClassificationModel
    
    Raw prediction for each possible label. The meaning of a "raw" prediction may vary between algorithms, but it intuitively gives a measure of confidence in each possible label (where larger = more confident). This internal method is used to implement transform() and output rawPredictionCol.
    
    Specified by:
    
    predictRaw in class ClassificationModel<Vector,DecisionTreeClassificationModel>
    
    Parameters:
    
    features - (undocumented)
    
    Returns:
    
    vector where element i is the raw prediction for label i. This raw prediction may be any real number, where a larger value indicates greater confidence for that label.
  - copy
```
public DecisionTreeClassificationModel copy(ParamMap extra)
```
    Description copied from interface: Params
    
    Creates a copy of this instance with the same UID and some extra params. Subclasses should implement this method and set the return type properly. See defaultCopy().
    
    Specified by:
    
    copy in interface Params
    
    Specified by:
    
    copy in class Model<DecisionTreeClassificationModel>
    
    Parameters:
    
    extra - (undocumented)
    
    Returns:
    
    (undocumented)
  - toString
```
public String toString()
```
    Description copied from interface: DecisionTreeModel
    
    Summary of the model
    
    Specified by:
    
    toString in interface DecisionTreeModel
    
    Specified by:
    
    toString in interface Identifiable
    
    Overrides:
    
    toString in class Object
  - featureImportances
```
public Vector featureImportances()
```
  - write
```
public MLWriter write()
```
    Description copied from interface: MLWritable
    
    Returns an MLWriter instance for this ML instance.
    
    Specified by:
    
    write in interface MLWritable
    
    Returns:
    
    (undocumented)

Class DecisionTreeClassificationModel

Nested Class Summary

Nested classes/interfaces inherited from interface org.apache.spark.internal.Logging

Method Summary

Methods inherited from class org.apache.spark.ml.classification.ProbabilisticClassificationModel

Methods inherited from class org.apache.spark.ml.classification.ClassificationModel

Methods inherited from class org.apache.spark.ml.PredictionModel

Methods inherited from class org.apache.spark.ml.Model

Methods inherited from class org.apache.spark.ml.Transformer

Methods inherited from class org.apache.spark.ml.PipelineStage

Methods inherited from class Object

Methods inherited from interface org.apache.spark.ml.tree.DecisionTreeModel

Methods inherited from interface org.apache.spark.ml.tree.DecisionTreeClassifierParams

Methods inherited from interface org.apache.spark.ml.tree.DecisionTreeParams

Methods inherited from interface org.apache.spark.ml.param.shared.HasCheckpointInterval

Methods inherited from interface org.apache.spark.ml.param.shared.HasSeed

Methods inherited from interface org.apache.spark.ml.param.shared.HasWeightCol

Methods inherited from interface org.apache.spark.ml.tree.TreeClassifierParams

Methods inherited from interface org.apache.spark.ml.param.shared.HasLabelCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasFeaturesCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasPredictionCol

Methods inherited from interface org.apache.spark.ml.param.Params

Methods inherited from interface org.apache.spark.ml.param.shared.HasRawPredictionCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasProbabilityCol

Methods inherited from interface org.apache.spark.ml.param.shared.HasThresholds

Methods inherited from interface org.apache.spark.ml.util.MLWritable

Methods inherited from interface org.apache.spark.internal.Logging

Method Detail

read

load

impurity

leafCol

maxDepth

maxBins

minInstancesPerNode

minWeightFractionPerNode

minInfoGain

maxMemoryInMB

cacheNodeIds

weightCol

seed

checkpointInterval

depth

uid

rootNode

numFeatures

numClasses

predict

transformSchema

transform

predictRaw

copy

toString

featureImportances

write