Skip to content

Functions

Functions let you wrap reusable logic into a single named block. You define a function once and call it many times. This keeps your code organized and reduces copy/paste errors.

Use the def keyword followed by a name and parameters. The return statement sends a value back to the caller.

def calculate_gc_content(sequence):
"""Calculates the GC content of a DNA sequence.
Args:
sequence: A string of DNA bases.
Returns:
The fraction of G and C bases in the sequence.
"""
sequence = sequence.upper()
gc_count = sequence.count("G") + sequence.count("C")
return gc_count / len(sequence)
print(calculate_gc_content("ATGCGATCGA"))
0.5
print(calculate_gc_content("AAATTTAAATTT"))
0.0

The triple-quoted string right after def is called a docstring. It describes what the function does. We use the Google style for docstrings throughout this project. More on that below.

You can give parameters default values. The caller can override them or leave them as is.

def label_significance(pvalue, log2fc, alpha=0.05, fc_threshold=1.0):
"""Labels a gene as up, down, or not significant.
Args:
pvalue: The p-value from a statistical test.
log2fc: The log2 fold change.
alpha: Significance threshold. Defaults to 0.05.
fc_threshold: Minimum absolute fold change. Defaults to 1.0.
Returns:
A string label: "up", "down", or "ns".
"""
if pvalue < alpha and log2fc > fc_threshold:
return "up"
elif pvalue < alpha and log2fc < -fc_threshold:
return "down"
else:
return "ns"
print(label_significance(0.001, 2.5))
up
print(label_significance(0.001, -1.8))
down
print(label_significance(0.5, 3.0))
ns

You can override defaults by name. This makes function calls easier to read.

print(label_significance(0.001, 0.5, fc_threshold=0.25))
up

A function can return several values at once using a tuple. You unpack them into separate variables on the calling side.

def summarize_counts(counts):
"""Summarizes a list of gene counts.
Args:
counts: A list of integer read counts.
Returns:
A tuple of (total, mean, nonzero_count, zero_fraction).
"""
total = sum(counts)
mean_val = total / len(counts)
nonzero = sum(1 for c in counts if c > 0)
zero_fraction = sum(1 for c in counts if c == 0) / len(counts)
return total, mean_val, nonzero, zero_fraction
gene_counts = [0, 15, 230, 0, 45, 120, 0, 78]
total, mean_val, nonzero, zero_frac = summarize_counts(gene_counts)
print(f"Total: {total}")
print(f"Mean: {mean_val}")
print(f"Nonzero genes: {nonzero}")
print(f"Zero fraction: {zero_frac}")
Total: 488
Mean: 61.0
Nonzero genes: 5
Zero fraction: 0.375

Variables created inside a function are local. They do not affect variables with the same name outside the function.

threshold = 0.05
def check_significance(pvalue):
"""Compare a p value with the local threshold.
Args:
pvalue: P value to assess.
Returns:
True when the p value is below the local threshold.
"""
threshold = 0.01
return pvalue < threshold
print(check_significance(0.03))
print(f"Global threshold: {threshold}")
False
Global threshold: 0.05

The function uses its own local threshold of 0.01. The global threshold stays at 0.05.

That example is the reassuring half, and read on its own it teaches the wrong lesson. Assigning a new value to a name inside a function is local. Mutating an object you were handed is not, because the caller’s name and the parameter point at the same object.

def add_flag(genes):
"""Append a marker to a gene list.
Args:
genes: Mutable list of gene names.
Returns:
The same list after mutation.
"""
genes.append("FLAGGED") # mutates the caller's list
return genes
my_genes = ["TP53", "BRCA1"]
add_flag(my_genes)
print(my_genes)
caller's list after add_flag: ['TP53', 'BRCA1', 'FLAGGED']

Rebinding instead leaves the caller alone:

def rebind(genes):
"""Return a copy of a gene list with a marker.
Args:
genes: Input list of gene names.
Returns:
New list containing the marker.
"""
genes = genes + ["REBOUND"] # builds a NEW list
return genes
caller's list after rebind: ['TP53', 'BRCA1']

Same-looking function, opposite effect on the caller. Whenever you read a function that takes a list, dict, DataFrame or array, check whether it mutates the argument.

A default is evaluated once, when the function is defined, so every call that does not pass one shares the same object.

def collect(gene, seen=[]):
"""Append a gene to the shared default list.
Args:
gene: Gene name to collect.
seen: List used to collect gene names.
Returns:
List containing collected gene names.
"""
seen.append(gene)
return seen
collect('TP53') -> ['TP53']
collect('EGFR') -> ['TP53', 'EGFR'] <- TP53 is still there

The fix is seen=None, then build the list inside:

def collect_safe(gene, seen=None):
"""Append a gene without sharing a default list.
Args:
gene: Gene name to collect.
seen: Optional list of existing gene names.
Returns:
List containing collected gene names.
"""
if seen is None:
seen = []
seen.append(gene)
return seen
safe: ['TP53'] then ['EGFR']
CUTOFF = 0.05
def is_significant(p):
"""Compare a p value with the configured cutoff.
Args:
p: P value to assess.
Returns:
True when the p value is below the cutoff.
"""
return p < CUTOFF # silent dependency

This works. It also fails in a fresh session, in a different script, or when someone changes CUTOFF three hundred lines away. A function that reads a global it was not given is the most common reason code that worked yesterday stops working today. Pass it in with a default instead.

global makes rebinding escape too, and is worth recognising in code you did not write:

counter = 0
def bump():
"""Increment the module level counter.
"""
global counter
counter += 1
counter after two bump() calls: 2

A lambda is a small anonymous function written on one line. It is useful for short operations you only need once.

genes = ["TP53", "BRCA1", "EGFR", "KRAS", "MYC"]
sorted_by_length = sorted(genes, key=lambda g: len(g))
print(sorted_by_length)
['MYC', 'TP53', 'EGFR', 'KRAS', 'BRCA1']

Use lambdas for simple sorting or filtering. For anything longer than one expression, write a regular function instead.

map applies a function to every item in a list. filter keeps only items that pass a test.

import math
counts = [100, 250, 50, 300, 75]
log_counts = list(map(lambda c: round(math.log2(c + 1), 2), counts))
print(log_counts)
[6.66, 7.97, 5.67, 8.23, 6.25]
pvalues = [0.001, 0.23, 0.04, 0.87, 0.005]
significant = list(filter(lambda p: p < 0.05, pvalues))
print(significant)
[0.001, 0.04, 0.005]

Both map and filter return iterators. Wrap them in list() to see the results. You can also achieve the same results with list comprehensions, which many Python programmers prefer for readability.

This project uses Google-style docstrings as the standard format. Every function you write should include one. Google-style docstrings use Args:, Returns:, and Raises: sections with indented descriptions beneath each heading.

This format works with documentation generators like Sphinx and its Napoleon extension. These tools can parse your docstrings and produce professional HTML documentation automatically. Writing good docstrings is a habit that pays off when your codebase grows.

Concept Syntax Example
Define a function def name(params): def gc(seq):
Default argument param=value alpha=0.05
Return a value return value return gc_count / length
Return multiple values return a, b, c return total, mean, zeros
Lambda lambda params: expr lambda g: len(g)
Map map(func, iterable) map(lambda c: c+1, counts)
Filter filter(func, iterable) filter(lambda p: p<0.05, pvals)
Docstring """...""" after def See examples above

Now that you can write your own functions, learn how to use code others have written. Head to Working with Packages to learn about importing and managing Python libraries.